The goal of lead optimization is
to improve the properties of a molecule such that it can become a drug or a
chemical probe. Early in a program, particularly one based on fragments, the focus
is often potency, but eventually drug metabolism and pharmacokinetics (DMPK) parameters
become important. I always enjoy reading optimization stories and seeing clever
use of design principles. But how effective is the design process? Or, to put it
another way, how often are improvements just luck? This is the subject of an
open-access paper just published in Nature by Bryan Roth, James Fraser,
Brian Shoichet, and collaborators at University of California San Franciso,
University of North Carolina Chapel Hill, and several other institutes.
The goal was to create a “random background”
against which to compare directed efforts. To do so, the researchers chose six
protein targets with which they had experience: the GPCRs α2AAR, µOR, and CB2;
the transporter SERT; and the soluble enzymes AmpC and Mac1. For each protein,
they then chose one to six ligands, 18
total, with IC50, KI, or EC50 values ranging
from 1 nM to 43 µM. These ‘parent ligands’ were then analyzed systematically to
see how they could be modified one atom at a time by adding a methyl group, a
halogen, or a phenol, or replacing an aromatic ring carbon with a nitrogen. A
decade ago, we wrote about the importance of synthetic tractability, and this
was a key consideration here. Ultimately the researchers made 257 molecules,
with each parent giving rise to between 5 and 31 analogs. These analogs were
then tested in a variety of assays.
Breaking things is usually easier
than improving them, and that turned out to be the case here: 30% of changes
caused a drop in activity of at least ten-fold. Perhaps more surprisingly, 29 of the 257 analogs (11.3%) had at least ten-fold better activity
than the parent molecules. Such increases were widespread: only one target (AmpC)
did not see a ten-fold increase, and 10 of the 18 parent molecules had more
potent analogs. In other words, the odds of improvement for any given chemical
series were better than even, even without using any structural insights from
the protein.
This latest work stands in good
company. Back in 2012 we highlighted an analysis of the effects of methyl substituents
across thousands of examples covering more than 100 proteins. That work came to
similar conclusions, with ten-fold improvements in 8% of cases. A more recent
paper looked at 633 published examples where added chlorine atoms improved
activity by at least ten-fold. This latest Nature paper finds methyl substituents
to be slightly more likely to give ten-fold improvements than chlorine atoms, while
adding fluorine was much less likely to yield such improvements, and swapping
an aromatic carbon for a nitrogen never did.
One might assume that gaining potency
by adding a methyl group or a chlorine atom is due solely to increased
lipophilicity, but this did not seem to be the case as assessed by calculating
LipE/LLE values.
Weaker ligands were slightly more
likely to be improved by making random changes, which makes sense if you
consider that a more potent molecule is likely to be highly complementary to
its binding site. Similarly, larger pockets were more likely to see ligand
activity improve with random changes.
The researchers
crystallographically characterized 22 analogs of varying affinities derived from
three parent ligands against the protein Mac1 to look for general lessons.
Although the analyses proved interesting, “many of the effects of the small
perturbation analogs could be explained post hoc from their structures; however,
fewer were easily anticipated.”
Of course, as noted above, affinity is just one
component of a chemical probe, let alone a drug, so the researchers also assessed
several in vitro DMPK parameters: microsome stability, plasma stability, plasma
protein binding, solubility, hERG inhibition, and membrane permeability.
Similar to the activity measurements, some analogs had improvements in one or
more properties, while others saw declines. Unfortunately, the analogs with
improved potency often had decreased values for in vitro DMPK parameters; none
of the 29 analogs with ten-fold improved potency had combined improvements for stability,
permeability, and plasma protein binding. As the researchers ruefully conclude,
“the changes in in vitro PK were, at best, orthogonal to affinity or potency
fold change, and most were, if anything, anti-correlated with it.” This “frustrated
landscape” is all too familiar to medicinal chemists.
Making and testing all these
molecules took a tremendous amount of work (the paper lists over 30 authors), and the researchers naturally
wondered whether computational methods would have saved them the effort. They
used free energy perturbation (specifically FEP+, from Schrödinger) to calculate
affinities and three additional programs to calculate in vitro DMPK parameters.
Although the affinity correlations were strong, there was still considerable variation
for individual compounds: of 34 analogs predicted to increase in binding energy
by more than 1 kcal/mol (about five-fold), only 14 actually did so, while 8
showed decreased activity. Meanwhile, while plasma protein binding and permeability
predictions correlated with experimental values, stability and solubility “were
essential uncorrelated.”
As of now, computers are still no
substitute for the good old-fashioned design, make, test cycle. This paper
suggests that, in addition to careful design, it may also be worth introducing
some randomness into your analogs, particularly when they are easy to make. Sound
advice indeed.
5 comments:
I see this study, Dan, as something of a stamp-collecting exercise and not of great value from the perspective of drug discovery scientists working on specific projects. When performing matched molecular pair analysis one should examine both the mean value for ΔpIC50 and the corresponding standard deviation. A relatively high percentage of matched molecular pairs exhibiting increases in potency greater than tenfold might also reflect higher variance in the ΔpIC50 values. While the article lists Hajduk & Sauer (2008) Statistical analysis of the effects of common chemical substituents on ligand potency JMC 51:553-564 DOI: 10.1021/jm070838y as reference 16, the citation in the text (Using tenfold improvement between analogue and parent as a benchmark for substantial impact, 10% of analogues meet this standard in the ChEMBL database [16]) makes no sense whatsoever.
Hi Pete,
Just to make sure I understand, are you arguing that some or most of the analogs with measured activity at least 10-fold better represent experimental error? If so, do the dose-response curves shown in Supplementary Figure 2 help ameliorate that concern?
As for stamp-collecting, I suppose you could say the same of the Protein Data Bank or Darwin's bird collection. Carefully curated collections can spur some powerful ideas.
What I’m getting at, Dan, is that one needs to be careful about coming to conclusions by comparing points in tails of distributions because what one sees may be due to differences in average values or to differences in variance. The observed variance in a ΔpIC50 distribution can be treated as a sum of the variance in the ‘true’ ΔpIC50 values and a term resulting from the uncertainty in the measurements.
Consider two compounds (A and B) that exhibit identical pIC50 values and suppose that it’s been shown that, in each case, the standard deviation for the replicate measurements is 0.3 (corresponding to a variance of 0.09). The variance in ΔpIC50 will be 0.18 which corresponds to a standard deviation of 0.4. This means that if we measure a single pIC50 value for each of A and for B we should expect (assuming Normal distribution) to pIC50_A exceed pIC50_B by 0.4 log units 16% of the time (and vice versa).
I would most definitely not dismiss observations made in the field by biologists or curation of existing data as stamp collecting. While I actually disagree with Rutherford’s assertion that “All science is either physics or stamp collecting", there is always a danger that running experiments for the sake of generating data can descend into something equivalent to philately. I have some worries about the OpenADMET initiative given that they’re invoking “ground truth” and the “Avoid-ome” blurs the distinction between pharmacodynamics and pharmacokinetics.
Hi Pete,
You seem to be proposing a theoretical argument about data quality and variation in general as opposed to looking at the dose-response curves actually provided, which look quite good to my eye.
As for running experiments for the sake of generating data, do your concerns extend to alanine scanning or random mutagenesis to understand protein function?
The point about assay variation is practical rather than theoretical, Dan, and my argument was that one needs to account for assay variation when doing matched molecular pair analysis by looking at the tails of distributions. Crappy concentration responses should ring alarm bells but you still need replicates in order to quantify assay variation.
I would not regard either alanine scanning or random mutagenesis to be stamp collecting provided that the experiments were being performed to gain understanding of specific proteins. Studying random proteins and presenting the results as being relevant to unrelated proteins would, in my view, constitute stamp collecting. Some general advice that I offer in NoLE (https://doi.org/10.1186/s13321-019-0330-2) is that “drug designers should not automatically assume that conclusions drawn from analysis of large, structurally-diverse data sets are necessarily relevant to the specific drug design projects on which they are working”.
Post a Comment