28 September 2026

Expanding the Scope of chemical space at Novartis

Without a good library, it is difficult to find a good hit, so library design is one of the most import aspects of a screening campaign. The size of a library also impacts how it can be screened. A virtual or DNA-encoded library might be billions (or even trillions) of compounds, but for most physical screens, such huge numbers would be unfeasible. Generally, fragment screens sample more chemical space but require more resources, while screens of larger molecules can be higher throughput. In a new (open-access) J. Chem. Inf. Model. paper, Johanna Jansen, Charles Wartchow, and colleagues at Novartis describe a middle ground, the “Scope Concept.”
 
The researchers, many of whom have fragment and biophysics experience, were interested in molecules that might be larger than the rule-of-three guidelines for fragments, but still small enough to sample reasonable swaths of chemical space. To this end they developed the Scope library. To be included, compounds need to have molecular weights between 200-400, measured or calculated solubilities of at least 100 µM, at least two moieties capable of forming hydrogen bonds, and a “Scope ring system,” meaning two rings connected by a 0-3 atom linker. Additionally, undesirable substructures were removed, and in-house compounds had to pass quality control. Originally compounds were required to have log P < 5, but the solubility limit weeded these out anyway. Some warhead-containing molecules were included, though the fraction is not disclosed.
 
The Novartis compound collection was queried using these criteria to choose a 10,000-compound core Scope set as well as a 40,000-compound extension set designed for higher throughput screens. While these numbers are not large in the context of high-throughput screening libraries, they are larger than most fragment libraries according to our polls.
 
The Scope libraries proved to be popular at Novartis: 18 screens were performed in the first year across three sites. Half of these were biochemical screens. Of the rest, six used SPR, two used DSF, and one used native mass spectrometry (nMS). Most of the biochemical screens used the full 50,000 library, while the biophysical screens used the 10,000 core set and sometimes expanded to the full set. For four of five targets screened using a biochemical assay, primary hit rates were similar or slightly higher for the Scope library than a diversity library, though the Scope library was screened at a 50-100 µM concentration as opposed to 10-50 µM.
 
Three targets are discussed in some detail. The first, trypanosomal CLK1 kinase, had a freakishly high primary hit rate of 21% when screened at 100 µM, compared to 2.5% for a diversity set screened at 12.5 µM. Another kinase had an 11% hit rate, while a third had a more reasonable 1.2% hit rate when the ATP-binding site was blocked, suggesting binding to this site. The researchers now screen kinases against lower concentrations of the Scope compounds due to the high ligandability of kinases in general.
 
The chemical structures of half a dozen hits are shown, with IC50 values ranging from 30 nM to 1.1 µM. Interestingly two of the molecules are covalent modifiers, and both are somewhat large for fragments, with 19 or 20 non-hydrogen atoms.
 
Another target screened was the molecular hub protein 14-3-3, which we wrote about most recently just a few months ago. A screen of the 10,000-compound Scope core set looking for stabilizers of 14-3-3σ and a peptide from the estrogen receptor yielded 86 hits, of which three were fully validated in orthogonal assays. One of these yielded a crystal structure and turned out to be a covalent modifier at a cysteine residue previously shown to be reactive. As with CLK1, the hit was on the larger side, with 22 non-hydrogen atoms. A separate publication by Colin Skepper and colleagues in ACS Med. Chem. Lett. describes this and three other covalent hits derived from diversity sets. Here too, the hits ranged from 20 to 24 non-hydrogen atoms.
 
The native mass spectrometry screen mentioned above was run against the main protease (MPro) of SARS-CoV-2. The 10,000 compound core set was screened in pools of four, with each compound at 25 µM. Of 61 primary hits, five validated with IC50 values < 25 µM. In contrast to many inhibitors of this protein, these were mostly non-covalent.
 
The Scope Concept is a nice addition to the field of library design. As we noted last year, covalent fragments may need to be larger than non-covalent fragments, and it is interesting that two of the three case studies identified covalent modifiers. I hope Novartis publishes a follow-up in a few years with lessons from additional screens.

21 September 2026

Design vs randomness in lead optimization across a frustrated landscape

The goal of lead optimization is to improve the properties of a molecule such that it can become a drug or a chemical probe. Early in a program, particularly one based on fragments, the focus is often potency, but eventually drug metabolism and pharmacokinetics (DMPK) parameters become important. I always enjoy reading optimization stories and seeing clever use of design principles. But how effective is the design process? Or, to put it another way, how often are improvements just luck? This is the subject of an open-access paper just published in Nature by Bryan Roth, James Fraser, Brian Shoichet, and collaborators at University of California San Franciso, University of North Carolina Chapel Hill, and several other institutes.
 
The goal was to create a “random background” against which to compare directed efforts. To do so, the researchers chose six protein targets with which they had experience: the GPCRs α2AAR, µOR, and CB2; the transporter SERT; and the soluble enzymes AmpC and Mac1. For each protein, they then chose one to six  ligands, 18 total, with IC50, KI, or EC50 values ranging from 1 nM to 43 µM. These ‘parent ligands’ were then analyzed systematically to see how they could be modified one atom at a time by adding a methyl group, a halogen, or a phenol, or replacing an aromatic ring carbon with a nitrogen. A decade ago, we wrote about the importance of synthetic tractability, and this was a key consideration here. Ultimately the researchers made 257 molecules, with each parent giving rise to between 5 and 31 analogs. These analogs were then tested in a variety of assays.
 
Breaking things is usually easier than improving them, and that turned out to be the case here: 30% of changes caused a drop in activity of at least ten-fold. Perhaps more surprisingly, 29 of the 257 analogs (11.3%) had at least ten-fold better activity than the parent molecules. Such increases were widespread: only one target (AmpC) did not see a ten-fold increase, and 10 of the 18 parent molecules had more potent analogs. In other words, the odds of improvement for any given chemical series were better than even, even without using any structural insights from the protein.
 
This latest work stands in good company. Back in 2012 we highlighted an analysis of the effects of methyl substituents across thousands of examples covering more than 100 proteins. That work came to similar conclusions, with ten-fold improvements in 8% of cases. A more recent paper looked at 633 published examples where added chlorine atoms improved activity by at least ten-fold. This latest Nature paper finds methyl substituents to be slightly more likely to give ten-fold improvements than chlorine atoms, while adding fluorine was much less likely to yield such improvements, and swapping an aromatic carbon for a nitrogen never did.
 
One might assume that gaining potency by adding a methyl group or a chlorine atom is due solely to increased lipophilicity, but this did not seem to be the case as assessed by calculating LipE/LLE values.
 
Weaker ligands were slightly more likely to be improved by making random changes, which makes sense if you consider that a more potent molecule is likely to be highly complementary to its binding site. Similarly, larger pockets were more likely to see ligand activity improve with random changes.
 
The researchers crystallographically characterized 22 analogs of varying affinities derived from three parent ligands against the protein Mac1 to look for general lessons. Although the analyses proved interesting, “many of the effects of the small perturbation analogs could be explained post hoc from their structures; however, fewer were easily anticipated.”
 
Of course, as noted above, affinity is just one component of a chemical probe, let alone a drug, so the researchers also assessed several in vitro DMPK parameters: microsome stability, plasma stability, plasma protein binding, solubility, hERG inhibition, and membrane permeability. Similar to the activity measurements, some analogs had improvements in one or more properties, while others saw declines. Unfortunately, the analogs with improved potency often had decreased values for in vitro DMPK parameters; none of the 29 analogs with ten-fold improved potency had combined improvements for stability, permeability, and plasma protein binding. As the researchers ruefully conclude, “the changes in in vitro PK were, at best, orthogonal to affinity or potency fold change, and most were, if anything, anti-correlated with it.” This “frustrated landscape” is all too familiar to medicinal chemists.
 
Making and testing all these molecules took a tremendous amount of work (the paper lists over 30 authors), and the researchers naturally wondered whether computational methods would have saved them the effort. They used free energy perturbation (specifically FEP+, from Schrödinger) to calculate affinities and three additional programs to calculate in vitro DMPK parameters. Although the affinity correlations were strong, there was still considerable variation for individual compounds: of 34 analogs predicted to increase in binding energy by more than 1 kcal/mol (about five-fold), only 14 actually did so, while 8 showed decreased activity. Meanwhile, while plasma protein binding and permeability predictions correlated with experimental values, stability and solubility “were essential uncorrelated.”
 
As of now, computers are still no substitute for the good old-fashioned design, make, test cycle. This paper suggests that, in addition to careful design, it may also be worth introducing some randomness into your analogs, particularly when they are easy to make. Sound advice indeed.

14 September 2026

Covalent allosteric inhibitors of NLRP3

NLRP3 has received considerable attention as an inflammatory disease target. The protein normally exists in an auto-inhibited conformation, but in response to various stimuli it changes shape to help form a multimeric complex called the inflammasome. This in turn activates various inflammatory pathways, including cell death. Several drugs targeting NLRP3 have gone into the clinic, but none have been approved. In a new open-access Br. J. Pharmacol. paper, Brian Cook, Gabriel Simon, and colleagues at Vividion describe a covalent inhibitor that binds to a previously unknown site.
 
As we wrote nearly a decade ago, Vividion pioneered the use of chemoproteomics to identify covalent binders in cells and cell lysates. While exploring the mechanism of the previously reported NLRP3 inhibitor MCC950, they found that cysteine 280 became resistant to modification by electrophiles in the presence of the ligand. MCC950 does not contain an electrophile, arguing against direct interaction with Cys280 and suggesting instead that binding causes a conformational change to the protein which makes Cys280 less accessible. In other words, protein inhibition correlated with Cys280 being less exposed.
 
A screen of Vividion’s library yielded no hits against Cys280 but more than 300 against cysteine 463. Interestingly, modification of C463 also caused a decrease in reactivity of Cys280, suggesting a conformational change similar to the one caused by MCC950. Thus, the researchers pursued Cys463 binders.
 
Two parameters were used to guide compound optimization, target engagement and promiscuity. The first number, TE50, measures the extent of covalent engagement of a given cysteine after 1 hour in cell lysates. The second, P500, measures the percentage of all measured cysteines in the sample labeled by at least 40% after incubation with 500 µM compound for 1 hour. One of the initial hits, VVD-124476, gave a low micromolar value for TE50 but labeled some 20% of the measured proteome. Indeed, it was previously reported as a covalent ligand for a completely different protein. To be useful as a NLRP3 inhibitor, the researchers needed to reduce promiscuity, and the most straightforward approach is to reduce general reactivity.
 

To decrease reactivity, the researchers swapped the acrylamide warhead with a butynamide (VVD-142397). Butynamides are reported to be less reactive than acrylamides, a result corroborated at Vividion with chemoproteomic data on 400 matched molecular pairs. The acrylamide to butynamide swap improved selectivity some 20-fold, though it also decreased potency against NLRP3. Increasing the core ring size regained activity, and rigidifying the molecule led to further improvements. Further optimization ultimately led to VVD-338213 and several related molecules.
 
Although the researchers do not report kinact/KI values, a back of the envelope calculation puts it around 1200 M-1s-1 for VVD-338213, a potency where cell activity might be expected. Happily this turned out to be the case: the molecule blocked IL-1β secretion from human cells at lower concentrations than MCC950. Importantly, in cells where Cys463 had been mutated to alanine, MCC950 was still active while the covalent molecules were inactive.
 
In addition to promising cell activity, many of the covalent molecules had reasonable stability in whole blood, good permeability, and low efflux. (I do wish the ligand efficiency reactivities of the molecules were reported; after all Vividion, introduced this metric.) Some of the molecules also showed good brain penetration. In vivo pharmacodynamic studies are complicated by the fact that Cys463 is not conserved in mice, but fortunately the researchers could run studies on transgenic mice in which human NLRP3 had been introduced. In these animals, IL-1β was decreased by both VVD-338213 and by MCC950.
 
The researchers were also able to obtain a cryo-EM structure of one of their molecules bound to NLRP3, which confirmed covalent binding to Cys463. More importantly, it revealed that the binding site is in a cryptic pocket some 20 Å from where MCC950 binds. (One wonders if the computational technique we discussed last week would have been able to predict this pocket.) Similarly to the WRN inhibitor Vividion has taken into the clinic, the molecule appears to make only hydrophobic contacts with the protein, with no polar interactions. Mechanistically, the “Cys463 ligands act as molecular doorstops, stabilizing the inactive conformation and preventing the structural rearrangements necessary for transition to the active inflammasome disc.”
 
This is a lovely paper and a case study in transforming a highly reactive molecule into a selective molecule with in vivo activity. It is interesting that the starting molecule, with 28 non-hydrogen atoms, flagrantly violates the rule of three. Last year we asked whether covalent fragments need to be larger than their non-covalent counterparts, and this paper provides another example where the answer is yes. Whether or not this chemical series ultimately advances, it is nice to have another chemical probe for this target.

08 September 2026

Computational prediction of cryptic pockets

Last week we highlighted how NMR-based fragment screening was able to identify cryptic pockets on a small protein. Experimental methods are (still) the gold standard, but they take considerable effort and resources. In a new open-access J. Am. Chem. Soc. paper, Neha Vithani, David LeBard, and collaborators at OpenEye, Bayer, Genentech, and Khartis provide a computational approach.
 
The researchers defined six categories of cryptic pockets: 1) side-chain motion, 2) loop motion, 3) secondary-structure motion, 4) secondary-structure change, 5) interdomain motion, and 6) pre-existing deep pockets with no access to the surface of the protein. They collected a set of 21 proteins with 26 cryptic pockets overall (some proteins contain more than one) whose apo (unbound) and holo (ligand-bound) structures were available in the Protein Data Bank. This “Encore Dataset” contains all six types of cryptic pockets, though some are represented by only two or three examples.
 
To computationally identify cryptic pockets, the researchers used weighted ensemble molecular dynamics (WEMD) simulations on the apo structures. The details are quite technical but involve focusing on “normal modes” of protein movement. To find pockets, xenon is included in the simulation; these large atoms will fill pockets that open up during the course of the simulation, so detecting bound xenon reveals pockets.
 
WEMD successfully identified 24 of the 26 cryptic pockets in the Encore Dataset. Interestingly, both failures were for the same protein, hepatitis C polymerase, one a preexisting deep pocket and the other a secondary-structure change from order to disorder. (This latter category is exemplified in one of the cryptic pockets we discussed last week, in which crystallography revealed that a helix became disordered upon ligand binding.) The researchers hypothesize that these types of order-to-disorder transitions may be particularly difficult to detect computationally because the models are designed to identify secondary structures, thus overstabilizing them.
 
So WEMD is able to correctly predict most cryptic pockets in the Encore Dataset, but does it also predict cryptic pockets that don’t exist? Paul Samulsen famously quipped that economists had “predicted nine out of the last five recessions,” which actually looks pretty good compared to WEMD, which predicted more than 30 potential cryptic pockets on the protein biotin carboxylase. For comparison, the software program fpocket, which we wrote about last year, predicts 13 for this protein.
 
To rank which pockets actually exist, the researchers turned to another computational technique called Target X, which was described in this Research Square article earlier this year. This proved reasonably successful at ranking experimentally observed pockets near the top, with the top two predicted biotin carboxylase pockets corresponding to known pockets.
 
This is an interesting paper, but I am left with more questions. First, while the Encore Dataset is nice, why didn’t the researchers (also) use CryptoSite, a previously published set of 93 proteins with cryptic sites, which we discussed in 2024? More importantly, how can drug hunters decide whether to target a specific cryptic pocket? For instance, does WEMD and Target X allow one to assess whether cryptic pockets can support high-affinity ligands? These would be interesting questions for a follow-up paper.