Showing posts with label chemical space. Show all posts
Showing posts with label chemical space. Show all posts

14 June 2026

Vast fields of biologically active chemical space

Exactly 17 years ago, I highlighted a paper from Brian Shoichet’s lab in which he hypothesized that only tiny regions of chemical space are biologically relevant. If so, Practical Fragments quoted, “a major reason why the screening of synthetic compounds ever finds notable hits is that our libraries are biased toward the sort of molecules that proteins have evolved to recognize.” The best scientists continually question their assumptions, and this is exactly what Brian, Bryan Roth, and collaborators at University of California San Francisco and University of North Carolina at Chapel Hill have done in a new open-access J. Med. Chem. paper.
 
The idea that molecules more closely resembling metabolites and natural products would be more likely to be biologically active is reasonable given the promiscuity of many naturally occurring molecules. As the researchers point out, dopamine signals through 5 receptors, while serotonin binds to 14. Nature is the ultimate recycler, constantly reusing chemical motifs.
 
But since 2009 there has been a massive growth of make-on-demand libraries, and these have increasingly diverged from known molecules. Brian has been among the most prolific explorers of this new chemical space, and as he noted in his keynote at DDC 2026, these gargantuan libraries are yielding more hits against more targets.
 
Are these hits more selective? In other words, are molecules that are less “bio-like” likely to bind to fewer targets? This is the question the new paper addresses, focusing on the 5-HT2A serotonin receptor (5-HT2AR), the target for a variety of drugs including psychedelics. The researchers have extensively studied this GPCR; we wrote about a successful computational screen here. In the new paper, the researchers compared hits from a previous “make-on-demand” (MoD) library screen of 1.6 billion molecules with a new screen of 3.5 million “in-stock” molecules. As predicted, the in-stock compounds were much more bio-like than the MoD compounds as assessed by several metrics (Tanimoto similarity, Avalon fingerprints, and RDK7 fingerprints, for the cheminformatics aficionados among you).
 
Both libraries were computationally screened using DOCK3.8 against 5-HT2AR. Of the top scoring hits, 85 molecules from the in-stock library and 384 molecules from the MoD library were tested in a radio-ligand displacement assay. Hit rates were nearly identical at around 24% each. Functional assays revealed the hits to be active, with some acting as agonists while others were antagonists.
 
When tested against related GPCRs, specifically 5-HT2BR and 5-HT2CR, the in-stock molecules were more promiscuous than the MoD molecules. But when tested against a panel of 318 GPCRs there were no differences: the in-stock molecules bound on average 1.7% of the receptors while the MoD molecules bound on average 2.1%. The similar promiscuity persisted even when focusing on the 55 aminergic GPCRs (5-HT2AR is an aminergic GPCR). In fact, the most promiscuous compound of all came from the MoD set and is distinctly not bio-like.
 
Of course, this is just one study, but I suspect it will generalize. In 2009 I questioned whether biologically relevant chemical space is really so sparse. As I wrote then, “consider a vast field of some crop that can only be harvested by night. There are lights scattered haphazardly throughout the field. One might expect that the crops immediately under the lamp posts would be harvested more intensively than crops in darker parts of the field, even if other areas are equally productive. In this scenario, the lamp posts reveal natural products and similar molecules, but much – or even most – of (unlit) chemical space may also be biologically active, it just hasn’t been sampled yet.”
 
There is grandeur in this view of chemical space, which is supported by the new paper. Imagine yourself in the midst of the field on a moonless night. Turn on your headlamp and step forward in any direction. Drug leads glimmer as far as your light can reach.

20 April 2026

Twenty-First Annual Fragment-Based Drug Discovery Meeting

Last week some 875 people attended the CHI Drug Discovery Chemistry (DDC) meeting in San Diego. I can’t do justice to the 40 or so presentations I attended over four days but can highlight some of the main themes.
 
Reversible fragments
Membrane targets such as G protein-coupled receptors (GPCRs) pose a challenge for biophysical methods, but three talks presented progress. Matthew Eddy (University of Florida Gainesville) described high-resolution magic angle spinning (HRMAS) NMR, which entails spinning isolated cellular membranes containing GPCRs at high speed (4 kHz!), which miraculously yields sharp NMR signals for bound ligands. Matthew demonstrated applications with the human adenosine A2A receptor and weak (mM) ligands. He noted that the technique can work with native, poorly expressed proteins, though data collection times can be upwards of 30 minutes.
 
Kris Borzilleri described using 19F NMR to find ligands against an orphan GPCR at Pfizer. The 2287 fragments screened yielded 87 hits, of which 38 confirmed by SPR. SAR studies eventually yielded low micromolar ligands, but these were difficult to advance in the absence of structure (see here for a more successful example from Merck).
 
Vanessa Porkolab (Eurofins Cerep) described using the Nanotemper Spectral Shift technology to screen 826 fragments against the adenosine A2A receptor at 300 µM, with a 9.2% hit rate. Many of these ligands stabilized the GPCR in a thermal shift assay and seven were even active (as antagonists) in a cellular assay.
 
Turning to soluble proteins, Paola Di Lello presented a case study from Genentech and Vernalis applying ligand-observed NMR to the protein phosphatase PTPN22. Subsequent protein-observed NMR revealed that most of the 16 validated hits bound to two pockets some distance from the active site. The fragments were optimized to mid-micromolar affinity but showed no functional activity.
 
And Charlotte Hodson presented the eIF4E story from Astex. As we discussed last year, this yielded a low nanomolar ligand that did not have the desired cellular effects. Charlotte noted that subsequent genetic experiments were consistent with the limited efficacy. Still, the target was sufficiently interesting that a chemical probe would have been pursued even knowing it would be high-risk.
 
Covalent ligands
Covalent approaches made appearances throughout the conference. Keriann Backus (UCLA) described chemoproteomic approaches to find cysteine-targeting ligands; she noted that gain of cysteine residues (such as G12C in KRAS) are the most common missense variants in cancer. Keriann also warned how covalent compounds can cause potentially misleading effects in cells, as she described in Nat. Chem. Biol. last year.
 
In 2021 we wrote about the SpotXplorer fragment library from György Keserű (Hungarian Research Centre for Natural Sciences). György has now prepared a PhotoXplorer library, which uses diazirine tags for photochemical screening, which we described here. The new library has produced high hit rates across a variety of targets. György also described a new sulfozone-based photoprobe that is easier to prepare than diazirines.
 
Kelly Craft recounted a DNA-encoded library (DEL) screen at AbbVie against the target BCL2A1, also known as BFL1. This produced an aldehyde-containing low micromolar binder that formed an imine with buried lysine 102. Uncomfortable progressing an aldehyde, the researchers sought to covalently engage cysteine 55, the same cysteine targeted by AstraZeneca, as we wrote about here. The progression included at least one dual-warhead molecule which was crystallographically confirmed to bind both the lysine and cysteine. The effort ultimately yielded cysteine-selective leads.
 
Earlier this year I described the dDRTC method we developed at Frontier Medicines for determining kinact/KI, and Svetlana Kholodar presented a nice overview of its scope and utility. My colleague Johannes Hermann spoke in more detail about our covalent technologies, particularly those using AI.
 
Chemical space and the exploration thereof
Brian Shoichet (UCSF) gave an entertaining and wide-ranging account of “directed and random walks in chemical space.” Brian has consistently been on the bleeding edge of high-throughput in silico screens, from 67,000 compounds in 2009 to 138 million molecules in 2019 to 4 billion molecules today. When docking artifacts are avoided (as we discussed here), bigger libraries consistently produce more potent hits for more targets – an observation strikingly consistent with Alex Shaginian’s in 2023 as HitGen expanded their DEL libraries from billions to more than a trillion molecules. Brian is developing methods to computationally screen the >4 trillion make-on-demand molecules now available from companies such as Enamine.
 
Direct-to-biology (DTB) approaches, which rely on microscale chemical reactions screened without purification, have become increasingly popular methods for exploring chemical space. Jack Sadowsky correctly stated that Carmot was the first company formed around this approach; we previously wrote about the role Chemotype Evolution played in the discovery of sotorasib. Jack described how Kimia, which spun out of Carmot, has continued to advance the technology, applying it to find inhibitors selective for single members of closely related kinase families.
 
Allan Jordan described how Sygnature Discovery is applying DTB in a variety of assays including microsome stability and crystallography. (We wrote about crude reaction screening by crystallography earlier this year.) Expanding beyond DTB, Allan called their platform direct-to-discovery, and discussed how it led to a preclinical candidate with STORM Therapeutics in just 18 months.
 
WuXi Apptec is also using DTB. Peichuan Zhang described starting with ligands derived from fragment and DEL screens against the E3 ligase GID4 to make PROTACs to degrade BRD4; DTB was used to explore a wide range of different linkers. And Daniel Blair (St. Jude) described using DTB and affinity selection-mass spectrometry (AS-MS) to find new molecular glues for the oncology target LCK.
 
Computers, DEL, and DTB are not the only way to explore chemical space. Last year we covered Tom Kodadek’s bead-based screening approach at University of Florida Scripps, and Tom presented two talks on the topic, one using macrocycles to find binders to difficult targets such as PTP1B and one using small molecules to find molecular glues.
 
Speaking of PROTACs and glues, plenary keynote speaker Alessio Ciulli (University of Dundee) discussed the “evolution and future of targeted protein degradation.” Alessio noted that there are >25 PROTAC degraders and >10 glues in the clinic, though these collectively target only a small number of E3 ligands, so there is plenty of opportunity for the area to expand.
 
For many of us in industry, drugs represent the most privileged points in chemical space, and these often look quite different than we assume, as Dean Brown (Jnana) noted in his recent analysis of 104 oral small molecule drugs approved by the FDA from 2020 to 2024 (which we mentioned here). Some drugs contain eye-raising moieties such as acetylenes, styrenes, N-O bonds, and nitro groups. Indeed, it is worth remembering that venetoclax, arguably the most successful fragment-derived drug, sports a nitro group.
 
But before getting too complacent, Jonathan Baell (Manas) warned about frequent hitters in libraries of FDA-approved drugs. He notes in Eur. J. Med. Chem. earlier this year that many commercial libraries are actually enriched for molecules that cause spurious biological activity. Jonathan calls on library vendors to remove particularly egregious compounds, though I’d settle for world peace.
 
I’ll close on that pleasant thought, but please feel free to comment. I hope to see you in San Diego next year April 19-22 for the twenty-second iteration of DDC.

02 June 2025

Small and simple, but novel and potent

Back in 2012 we wrote about GDB-17, a database of possible small molecules having up to 17 carbon, oxygen, nitrogen, sulfur, and halogen atoms, most of which have never been synthesized. Although novelty isn’t strictly necessary for fragments, as evidenced by the fact that 7-azaindole has given rise to three approved drugs, it’s certainly nice to have. In a new (open-access) J. Med. Chem. paper, Jürg Gertsch, Jean-Louis Reymond, and colleagues at the University of Bern synthesize fragments that had not been previously made and show that they are biologically active.
 
When you start drawing all possible small molecules you get lots of weird stuff, including an explosion of compounds containing multiple three- and four-membered rings, which may be difficult to make. The researchers wisely focused on “mono- and bicyclic ring systems containing only five-, six-, or seven-membered rings.” They further limited their search to molecules containing just carbon and one or two nitrogen atoms (as well as hydrogen, of course). Systematic enumeration led to 1139 scaffolds, ignoring stereochemistry, of which 680 had not been previously reported in PubChem. Out of these, three related scaffolds were chosen for investigation.
 
Computational retrosynthesis was used to devise routes to the three bicyclic scaffolds, and these were successfully synthesized, along with mono-benzylated versions, for a total of 14 molecules (including stereoisomers), all rule-of-three compliant. The online Polypharmacology Browser 2 (PPB2) was used to predict targets, and several monoamine transporters came up as potential hits. The molecules were tested against norepinephrine transporter (NET), dopamine transporter (DAT), serotonin transporter (SERT), and the σ-R1 receptor in radioligand displacement assays. None of the free diamines were active, but several of the benzylated compounds were, in particular compound 1a.
 
Compound 1a was initially made as a racemic mixture, and when the two enantiomers were resolved (R,R)-1a was found to be a mid-nanomolar inhibitor of NET while (S,S)-1a was 26-fold weaker. Compound (R,R)-1a was also a mid- to high nanomolar inhibitor of σ-R1, DAT, and SERT. Pharmacokinetic experiments in mice revealed that the molecule had poor oral bioavailability but remarkably high brain penetration and caused sedation. The researchers conducted additional mechanistic studies beyond the scope of this blog post and conclude that (R,R)-1a could be a lead for “neuropsychiatric disorders associated with monoamine dysregulation.”
 
There are several nice lessons in this paper. First, as we noted more than a decade ago, there is plenty of novelty at the bottom of chemical space. Moreover, and in contrast to our post last week, even small fragments can have high affinities. But novelty comes at a cost: synthesis of compound 2a required eight steps from an inexpensive starting material with an overall yield of just 9%, though this could certainly be optimized. Nonetheless, particularly for CNS-targeting drugs which usually need to be small in order to cross the blood brain barrier, the price might be worth paying.
 
Of course, even within this paper there are hundreds more scaffolds to look at than the three tested, and perhaps the researchers were lucky that their choices were biologically active. As computational methods continue to advance, it will be worthwhile turning them loose on GDB-17.

20 March 2023

Versatile fragments from the Protein Data Bank

Four years ago we highlighted an analysis of fragments taken from the Protein Data Bank (PDB). Of 462 unique fragments, just 21 bound in more than one pocket. With the assumption that such “versatile” fragments may be particularly valuable starting points, Esther Kellenberger and colleagues at CNRS Univeristé de Strasbourg have done their own exploration of the PDB, as reported (open access) in Front. Chem.
 
Structures deposited in the PDB starting in 2000 with resolution better than 3 Å were examined to find those containing fragment-sized molecules (MW < 300 Da). Crystallization additives, phosphate and sulfate ions, and other unlovable molecules such as PAINS were excluded. Further triaging for fragments that bound in more than one pocket and in more than one binding mode (ie, different types of interactions) ultimately yielded a set of 203 versatile fragments. (One reason why so many more fragments were found in this study is the fact that the previous analysis required the word “fragment” to be present in the PDB entry.)
 
The versatile fragments are mostly compliant with the rule of three, with violations mostly related to the number of hydrogen bond donors or acceptors. Only a single molecule had ClogP > 3, though 50 were quite hydrophilic, with ClogP < 0. Interestingly, 45 of the molecules are listed as small molecule drugs, and 98 are substructures of approved drugs. Perhaps this is not surprising; drugs themselves are studied particularly intensively and frequently included in screening libraries.
 
The researchers had previously analyzed commercial libraries, and in the new paper they compared versatile fragments with the SpotXplorer library we wrote about here and the functionally diverse fragments used at XChem. Surprisingly there was very little overlap, even though most of the versatile fragments or analogs are commercially available. That said, some of the versatile fragments are molecules one may not want in a fragment library, such as the cofactor lipoic acid and the metal chelator 1,10-phenanthroline.
 
Binding modes for the same fragment in different pockets could vary considerably. The “universal fragment” 4-bromopyrazole, which we wrote about here, bound in two different binding modes, while the nucleoside thymidine showed a whopping 26 different binding modes. Conformations of the fragments could vary too, with only 43% of fragments showing a conserved conformation in all binding sites (defined as < 0.5 Å RMSD). Conformational changes, along with different protonation states, could be among the reasons why predicting fragment binding continues to be challenging.
 
This is a nice analysis, and it may be worth adding some of these versatile fragments to your own library. Laudably, SMILES strings for of all of them are provided in the supplementary material.

13 March 2023

A very useful list: common linkers and bioisosteric replacements

Last week’s post highlighted an example of fragment linking, which despite being less common than fragment growing can still be effective. But how do you choose the linker? We’ve previously written about the most common rings found in drugs. In a new Bioorg. Med. Chem. paper Peter Ertl and colleagues at Novartis tabulate the most common linkers found in bioactive molecules.
 
The researchers start by defining linkers “as moieties connecting 2 ring systems.” To focus on druglike molecules, linkers could contain no more than eight non-hydrogen atoms total and no more than five consecutive bonds between the two ring systems. This means that para-disubstituted phenyl or 1,4-disubstiuted butyl would both be considered in the analysis, but longer linkers such as this recent example would not.
 
Molecules were extracted from the databases ChEMBL and ZINC, yielding a total of 1686 unique linkers. Various descriptors were calculated for all, which in addition to size and length included the number of heteroatoms and electronic properties. Bioactivity data for molecules in ChEMBL was used to assess which replacements were most frequently tolerated. If one linker could be replaced by another without causing a drop in affinity (or inhibition, etc.), the two linkers were considered to be bioisosteres.
 
So, what are the most common linkers? A single methylene is the most common, followed by an amide bond. I was surprised that, of the 40 most common linkers, only five are rings: para-disubstituted phenyl, 1,4-piperzine, 1,4-piperidine, 1,2,4-oxadiazole, and meta-phenyl, in that order. Not coincidentally, phenyl rings, piperidines, and piperazines are also the most common rings found in drugs, according to an analysis last year.
 
Last year we highlighted a paper from the Ertl group that included a link to a “Ring Replacement Recommender,” which suggests bioisosteric replacements for any ring. Alas, there is no “Linker Replacement Recommender,” but the new paper does provide a “bioisosteric replacement network,” which is a full-page 10 x 15 grid with the 150 most common linkers arranged such that nearby linkers are likely to be bioisosteric. For example, para-phenyl is adjacent to 2,5-thiophene and quite some distance from sulfone. These make sense, but there are also less obvious examples: the table suggests that a 1,4-pyrazole makes a good replacement for a carbamate.
 
The next time you’re doing SAR, it may be worth consulting the bioisosteric replacement network for ideas.

05 September 2022

Is phenotypic fragment screening worthwhile?

Fragment-based drug discovery is almost always target-based. Indeed, not until the development of powerful biophysical techniques such as protein-labeled NMR did FBLD really began in earnest. Phenotypic fragment screens against cells, tissues, or animals are uncommon. In an open-access Front. Pharmacol. paper, Chris Lipinski and Andrew Reaume (Melior Discovery) argue that they should be used more often.
 
The researchers analyzed all 184,139,678 compounds in the CAS registry with molecular weights between 100 and 999 Da. These were divided into 18 bins (100-149 Da, 150-199 Da, etc.) Next, they calculated the percentage of molecules within each bin with any biological data as evidenced by the “biological study” tag in SciFinder-n.
 
In terms of raw numbers, fragments are well-represented, with the 250-299 Da bin containing close to 40 million molecules. However, only about 4% of these had any biological data. Molecules with molecular weights between 300 and 549 were abundant and also had considerably more biological data – up to roughly 50% of compounds in the 500-549 Da bin. In other words, people don’t seem to be screening lower molecular weight compounds in biological assays as often as they are screening larger molecules.
 
The assumption may be that small fragments are not biologically active, but the researchers revisit a classic In the Pipeline post in which Derek Lowe lists 56 drugs with molecular weights equal to or less than that of aspirin (180 Da). Most of these are old drugs, with all but three first reported in the chemical literature before 1980.
 
The researchers suggest that more effort should go into exploring the biology of smaller molecules, particularly those for which some activity is already reported. They also draw an interesting distinction between two uses of the word pleiotropic. People often say that a drug has pleiotropic effects if it acts on multiple targets; a classic example is imatinib, which hits several kinases in addition to the target BCR-ABL. However, the term pleiotropic originates in genetics and initially referred to one gene having multiple effects. Thus, a drug that acts on a single protein can have multiple effects, as in the case of the PDE5 inhibitor sildenafil.
 
As an example of a pleiotropic fragment, the researchers discuss MLR-1023, a fragment-sized molecule first discovered in a phenotypic screen at Pfizer in the 1970s. The molecule has shown promise in disease models ranging from atherosclerosis to myeloproliferative neoplasms and was taken into the clinic by Melior in 2014 as an anti-diabetic agent. All of these varied effects seem to stem from the ability of the compound to act as an activator of Lyn kinase. With just 15 non-hydrogen atoms and a molecular weight of 202 Da MLR-1023 is comfortably within rule of three space. Despite its small size, the molecule is a potent activator of Lyn, with an EC50 around 50 nM, giving it a ligand efficiency of 0.66 kcal/mol per heavy atom.
 
Is MLR-1023 an outlier or an example of an underexplored pool of pharmacological riches? My suspicion is the former. It is rare to find fragments with EC50s < 1 µM, let alone < 100 nM. Moreover, I suspect that many proteins are so difficult to drug that a molecule will need to be well beyond fragment-space – and even rule-of-five space – to have an effect. The protein-protein interaction targeted by venetoclax (MW = 868 Da) immediately comes to mind.
 
That said, the idea that a large group of tiny molecules is underexploited is worth exploring. For some types of drugs perhaps we don’t need extreme potency: Mike Hann noted a decade ago that the EC50 values of approved drugs average 20-200 nM and cautioned against an “addiction to potency.” And because fragments are likely to have low affinities towards most proteins, they may even be more specific than larger drugs. It will be fun to discover how much room there really is at the bottom.

01 August 2022

What rings are found in drugs?

Recently we highlighted the “Ring Replacement Recommender,” which provides suggestions for how to improve affinity by replacing one ring with another. The recommendations are based on an analysis of hundreds of thousands of molecules. But what about the rings found in actual drugs? This is the focus of a J. Med. Chem. paper by Richard Taylor and collaborators at UCB and Bohicket Pharma Consulting.
 
The researchers examined FDA-approved and investigational drugs with disclosed structures as of January 2020. These were fragmented into component “ring systems” for analysis. (Ring systems include not just monocycles but fused rings, such as purine. For example, sotorasib consists of four ring systems: benzene, pyridine, piperazine, and pyrido[2,3-d]pyrimidin-2-one.) More than 90% of drugs contain at least one ring.
 
Approved drugs have just 378 unique ring systems in total – a small increase from when the researchers examined approved drugs in 2014. The phenyl ring is found 727 times, with pyridyl (86 examples) a distant second, followed by piperidine (76 examples) piperazine (65 examples) and cyclohexane (47 examples). After that the numbers drop off sharply, with pyrazine in 50th place with just six examples and fluorene in 100th place with three examples.
 
Investigational drugs at first appear to be more diverse, with 450 unique ring systems, 280 of which are not found in approved drugs. Of these 280, pyridazine is the most common, with nine examples, followed by oxetane, with seven, but things quickly become less common from there, with 271 of the ring systems found just once. In contrast, ring systems found in drugs are found in multiple compounds, and in fact two thirds of investigational drugs only contain previously used ring systems.
 
Many of the new ring systems are closely related to those found in approved drugs, with nearly half differing by at most two atoms. Perhaps because of this the overall properties of the ring systems are similar between approved and investigational drugs, with no significant differences in heteroatom ratio, percentage of sp3 centers, or number of rings per system.
 
What new opportunities exist? The researchers identified nearly half a million synthetically accessible ring systems and winnowed these down to 3902 ring systems that have similar heteroatom ratios to those found in drugs and differ by at most two atoms. This attempt to explore new chemical space is similar to earlier work from the same group (here) as well as that from others (here, here and here).
 
The researchers also examined growth vectors and combinations of rings, the latter by using graph theory. These analyses suggest that investigational drugs do have greater variety. In other words, even if the component rings are shared with approved drugs, they might be combined in new ways.
 
Whether certain ring systems are more likely to fail in the clinic was intentionally not addressed, due to the difficulty of assessing why the failures occurred. For example, drugs can fail for commercial reasons; a company may choose to drop a drug against a particular target rather than be tenth to market. And even when the failure is due to the science, it might not be an indictment of the drug itself. Verubecestat did lower β-amyloid levels in people as designed, but had no effect on Alzheimer’s disease.
 
This paper is a fun read, and it will likely provide ideas for scaffold hopping and library design. It is also a reminder of how much chemical space remains to be explored.

06 June 2022

What to make first? A new “Ring Replacement Recommender” provides suggestions

So you’ve run a fragment screen, gotten some hits, and validated them. What then? Looking for in-house or commercial analogs is always a good idea, but if you’re serious about a project you’ll eventually need to do chemistry, for example replacing one ring with another (say, a pyridyl for a phenyl). The possibilities are almost endless, especially if you don’t know how your fragment binds. In a new Eur. J. Med. Chem. paper, Peter Ertl and colleagues at Novartis describe a “Ring Replacement Recommender” to rapidly improve biological activity.
 
To determine which replacements are likely to improve affinity, the researchers turned to ChEMBL, a database of more than 2 million molecules and associated biological activity extracted from tens of thousands of publications. From these, more than 68,000 chemical series were chosen for analysis. Each series had on average 16 members, and at least three. The biological activity of each member of a series was compared with other members of the same series. (Importantly, the researchers intentionally excluded anti-targets such as hERG and CYPs so the tool wouldn’t inadvertently improve binding to these.) Focusing only on ring replacements that were reported in at least five publications led to a set of 26,762 changes. Changes could be as modest as adding a methyl substituent or more elaborate such as changing a single aromatic ring to a fused aromatic-aliphatic ring system.
 
One would think that most changes would have little effect, as had previously been seen in the case of methyl additions. Indeed about 65% of the replacements caused shifts in potency of 2-fold or less, which is probably within experimental error. However, 2860 replacements of 245 rings improved affinity at least 2-fold (averaging 3.5-fold), with 223 cases yielding greater than ten-fold improvements.
 
Analyzing the data further, the researchers found 80 ring systems that frequently led to improvements in affinity, and they suggest these could be used as “universal” or privileged building blocks. Strikingly, 74 of these are aromatic, confirming work from Cohen we highlighted in 2020 that proteins may favor “flat” rather than shapely molecules.
 
The researchers also extracted 9515 drugs and clinical compounds from ChEMBL and examined the component fragments. Of the 80 ring systems in the universal set, 19 are found in 50 or more drugs, with another 37 found in at least 5 drugs. This set may be a particularly attractive go-to list.
 
Importantly, not only are all the replacements available in the Supporting Information, the researchers have created a handy and free online tool. Just click on a ring of interest and the Ring Replacement Recommender provides suggestions, along with the average fold improvement observed and the number of publications used for the calculation.
 
To see how well it works, I looked at a couple recent examples which entailed ring changes. The indole to indazole replacement used in the TLR7/8 work described last month was not suggested by the Recommender, though in that case the researchers had the benefit of a crystal structure. On the other hand, a cyclobutyl to phenyl substitution for SARS-CoV-2-3CLp was correctly predicted to be beneficial.
 
Of course, as we’ve said repeatedly, affinity is only part of the battle in drug discovery, and the researchers emphasize that their recommendations may not improve physicochemical or pharmacokinetic properties. But for the earliest stage of a program, and especially in the absence of other data, it’s worth giving the Recommender a try.

31 January 2022

A framework for evaluating commercial fragment libraries

The easiest way to build a fragment library is to purchase one. Quite a few vendors sell fragments, and as our poll from a few years ago demonstrated, most buyers are quite happy with them. But what exactly do they offer? This is the subject of a new paper by Gilles Marcou, Esther Kellenberger, and colleagues at CNRS Université de Strasbourg in RSC Med. Chem.
 
The researchers analyzed 86 different libraries from 14 vendors that were available in February of 2021. These were classified into ten categories, such as “general,” “3D-shaped,” “metal chelating,” “diverse,” “covalent,” etc. Individual library sizes ranged considerably: 41 had ≤ 2000 compounds, 31 had 2000-10,000, and 14 libraries had > 10,000 molecules. The total number of fragments came to 754,646, of which 512,284 were unique, indicating some redundancy between libraries. Laudably, the structures and several analyses are all provided as downloadable files here.
 
Most of the fragments are 200-300 Da, with only 13% less than 200 Da. This skew towards larger molecules is common but may not be desirable, as researchers at Astex demonstrated back in 2014. On the other hand, people do seem to be paying attention to lipophilicity: nearly half the fragments have AlogP < 1. Interestingly, less than a quarter of fragments strictly fulfill the rule of three, though the majority of violations are for more than 3 hydrogen bond acceptors, which is probably not as important as the other criteria, according to analyses of approved drugs.
 
Two different methods were used to assess diversity. These were applied to 433,433 compounds from fifty libraries; specialized libraries such as fluorine-rich and covalent libraries were excluded. The first analysis deconstructed fragments into 59,270 component scaffolds. Not surprisingly, benzene was the most common, present in nearly 5% of all fragments. Quinoline, indole, pyridine, and benzimidazole all were present in at least 1% of compounds. At the other end of the spectrum, 36,555 scaffolds occurred only once. Not surprisingly, these tended to be more complex.
 
In addition to assessing scaffolds, the researchers developed a “Generative Topographical Map (GTM) model to represent the chemical space in a landscape.” The resulting figures do indeed look like topographical maps, with darker regions corresponding to more populated areas. For example, since substituted benzimidazoles are common and similar to one another, they form a dark cluster. Not unexpectedly, the landscape for the set of 433,433 compounds is heterogenous, with denser regions separated by sparsely-populated regions.
 
A nice feature of the GTM model is that it allows easy, intuitive comparisons. For example, some of the “diverse” libraries are more diverse than others, or emphasize different regions of chemical space, and potential customers may want to take these into account.
 
Fragment shapeliness was assessed using plane of best fit (PBF), where lower values correspond to “flatter” molecules, such as benzene, with PBF = 0. The libraries varied considerably in their average PBF, though reassuringly the “3D-shaped” libraries did have higher values. Interestingly, GTM models showed both flat (PBF < 0.1) and non-planar (PBF ≥ 0.1) fragments had similar distributions across fragment space.
 
Overall this is a valuable snapshot of the current state of commercial libraries, and makes a useful complement to the ongoing analysis Chris Swain does at Cambridge MedChem Consulting. Of course, the devil is in the details; PAINS still sometimes show up in commercial libraries, and quality control can vary. In the end you’ll want to do your own vetting, but this is a good place to start.

06 September 2021

How fragments become leads

Our most recent poll asked how often synthetic challenges had kept researchers from pursuing a particular fragment or had impeded a fragment-to-lead project. Around two-thirds replied sometimes or often. A new open-access paper in Chem. Sci. by Rachel Grainger, Rhian Holvey, and colleagues at Astex does a deep dive into fragment-to-lead chemistry, provides a powerful visual tool, and ends with something of a call to action.
 
The researchers take as their starting points 131 fragment-to-lead success stories published from 2015 to 2019 and collated in a series of five J. Med. Chem. Perspectives. All of these started with fragments (< 300 Da) for which affinity increased by at least 100-fold, with the resulting leads having affinities of 2 µM or better. As the new paper points out, this could introduce “survivorship bias,” in that less successful projects are not included. However, as the point of the paper is to figure out what works, this likely strengthens the conclusions.
 
The targets themselves are fairly diverse: 24% kinases, 9% proteases, 36% other enzymes, 11% bromodomains, 14% other protein-protein interactions, and 6% other types of targets. The researchers closely examined how the leads related to the initial fragments. Full details are provided in the Supplementary Material (pdf). The researchers have also constructed a handy interactive viewer you can use to do your own analyses. Here is an overlay of an initial fragment (taupe space-filling) with the final lead (yellow surface).
 

What are the results? The first observation is that 93% of leads have at least one polar interaction (such as a hydrogen bond) that is conserved from the initial fragment. The most common functional groups making direct contacts to proteins are N-H hydrogen-bond donors (35%) followed by aromatic nitrogen hydrogen-bond acceptors (23%) and carbonyl oxygen hydrogen-bond acceptors (22%).
 
The second observation is that over 80% of fragments are grown from one or two vectors (examples of one and two, with the second the subject of the figure above). This is perhaps not surprising; growing from three or more vectors would likely result in portly molecules that may be more difficult to advance, venetoclax notwithstanding.
 
But the really interesting observation is that the majority of growth vectors (~80%) originate from carbon atoms. Moreover, more than half of the bonds formed are carbon-carbon bonds. For the non-chemists in the audience, this is significant because carbon-carbon bond forming reactions are not always straightforward, particularly in the presence of polar moieties.
 
In the early days of FBLD, one hope was that including functional groups such as amides in a fragment collection would facilitate fragment growing. The new paper suggests that this is naïve: a functional group in a fragment is likely to interact with the protein and so block potential growth vectors. Indeed, only 18% of growth vectors come from N-H groups, despite the fact that these are among the most synthetically accessible.
 
These findings thus explain why fragment-to-lead efforts can be so challenging. The researchers provide an example of a chemical series they ultimately abandoned due to poor synthetic tractability.
 
The paper also builds on earlier papers from Astex exhorting chemists to further advance chemical methodology. As they conclude:
An “ideal synthesis” of a lead would allow: (1) site-selective formation of bonds at all growing points of a fragment, (2) whilst being mild enough to be compatible with essential polar functionality, and (3) proceeding with minimal or no need for protecting groups….
 
We believe that further development of C-H functionalisation that is tolerant to polar fragments has the potential to transform FBDD.
If you’re in academia, this looks like a good opening for a grant proposal!

07 June 2021

A minimal fragment library for maximal coverage of pharmacophore space

Last week we described a fragment library built with the aid of machine learning and designed to contain privileged fragments that should produce high hit rates. Unfortunately, only about a tenth of the library members are commercially available, so it will be some time before we know whether the design was successful. We continue the theme of fragment libraries with a just published Nat. Commun. paper by György Keserű (Hungarian Research Centre for Natural Sciences) and a large group of multinational collaborators (see also here for a nice summary by György).
 
The researchers started by analyzing more than 3300 crystal structures of protein-fragment complexes in the protein data bank. Fragments were defined as having 10-16 non-hydrogen atoms, and the computational approach FTMap was used to ensure that fragments were binding at hotspots as opposed to spurious, less ligandable sites. This exercise yielded 3584 fragments, but many of them were identical or very similar to one another. The researchers used a series of computational tools to cluster similar fragments (or pharmacophores) and choose a set that would maximize diversity. This ultimately led them to assemble a library of just 96 fragments, purchased from five vendors.
 
This SpotXplorer0 library mostly follows the rule of three, with 7 to 17 non-hydrogen atoms, MW 100-250 (or 280 for bromine-containing molecules), ≤ 3 hydrogen bond donors, ≤ 8 hydrogen bond acceptors, and ≤ 3 rotatable bonds. In addition, all members have 1-3 rings, no more than a single halogen or sulfur atom, and no PAINS. Despite the small size, this library covers most of the pharmacophores identified in the larger set, and considerably more than the F2X-Entry fragment library we highlighted last year or the top five commercial library vendors we noted here.
 
The researchers then screened this library against eight targets. Three GPCRs (the serotonin receptors 5-HT1A, 5-HT6, and 5-HT7) were assessed in a cell-based radioligand displacement assay with fragments at just 10 µM. Despite the low concentration, 4-11 hits were found. Biochemical screens conducted at 800 µM against the proteases thrombin and Factor Xa yielded 7 and 8 hits respectively. Further analysis revealed that the SpotXplorer0 ligands sampled a majority of the pharmacophores found in published fragment hits against theses five targets.
 
Next the researchers screened their library against the histone methyltransferase SETD2, an oncology target with few known attractive ligands. An enzymatic assay yielded two hits, with IC50 values between 300 and 500 µM.
 
Finally, the SpotXplorer0 library was part of the XChem crystallographic screens against the SARS-CoV-2 main protease (Mpro) and Nsp3 macrodomain, which we discussed here and here. For Mpro, just a single hit was found. This is only half the overall hit rate for noncovalent fragments in the crystallographic screen against this target, but the hit is functionally active and has a high ligand efficiency.
 
The screen against NSP3 yielded five hits binding at two different sites, for a hit rate of 5.2%. The overall hit rate against this target was 8%, but that encompasses screens against two crystal forms of the protein. The crystal form used for SpotXplorer0 had a hit rate of 21%.
 
In summary, SpotXplorer0 is new fragment library that gives high coverage of experimental pharmacophore space. Laudably, structures of all 96 fragments are provided in the Supplementary Information. But the jury remains out on how hit-rich the library will be. Interestingly, the F2X-Entry library we highlighted last year gave considerably higher hit rates of 21% and 30%, albeit against two different targets. SpotXplorer0 is being screened crystallographically against multiple targets at XChem, and it will be interesting to see how it performs in the long run.

17 October 2016

FBLD 2016

Last week the sixth FBLD meeting was held in Cambridge, MA. Like its predecessors in 2014, 2012, 2010, 2009, and 2008, this meeting was an enormous success, mixing more than 230 scientists with excellent (and liberal) food and drink. With 33 talks, more than 30 posters, and several vendor booths and workshops I won’t be able to do more than capture a few highlights.

The most striking feature for me was the number of success stories. This began with Steve Fesik’s keynote lecture, in which he discussed the MCL-1 inhibitors he and his team at Vanderbilt have discovered. When we highlighted his work last year he had reported low nanomolar inhibitors, but these did not have cell-based activity. His group has now optimized the molecules to low picomolar biochemical potency, low nanomolar cellular activity, and good activity in mouse xenograft models. This has not been easy: more than 2210 compounds were made, guided by 60 X-ray structures and dozens of pharmacokinetic experiments. It seems to be paying off though, and the researchers are developing biomarkers with the goal of advancing a compound into clinical testing.

Two other notable success stories about clinical candidates must be mentioned, though I’ll wait until publications come out before going into detail. Kathy Lee described how she and her colleagues at Pfizer chose a fragment that was less potent and ligand-efficient than other hits due to its interesting binding mode and were able to advance it to PF-06650833, an IRAK4 inhibitor with potential for inflammatory diseases. And Wolfgang Jahnke discussed how he and his colleagues at Novartis were able to discover and advance ABL001, an allosteric inhibitor of BCR-ABL, despite having the project halted twice – a reminder that persistence is essential.

Several other success stories have been covered at least in part on Practical Fragments, including inhibitors against PDE10A (presented by Izzat Raheem of Merck), Dengue RNA-dependent RNA polymerase (presented by Fumiaki Yokokawa of Novartis), lipoprotein-associated phospholipase A2 (presented by Phil Day of Astex), and BACE1 (presented by Doug Whittington of Amgen).

Crystallography was another theme, and several of the success stories relied on crystallographic fragment screening. Frank von Delft of the Structural Genomics Consortium described developments that allow screening 1000 crystals per week at Diamond’s Xchem facility in the UK, which include acoustic dispensing of compounds into crystallization drops – while carefully avoiding hitting the crystals head-on.

Several computational talks reported results that run contrary to conventional wisdom. Vickie Tsui of Genentech discussed their CBP bromodomain program (which we recently discussed here). Several water molecules form a highly ordered network in the protein, and a WaterMap analysis suggested that these were high-energy and that displacing them would lead to an enhancement in activity. Unfortunately this turned out not to be the case, though the researchers were able to get to low nanomolar inhibitors by growing towards a different region of the protein.

Li Xing mined the Pfizer database of 4000 kinase-ligand structures to extract 595 unique hinge binders. Not surprisingly, some of these – such as adenine and 7-azaindole – bound to multiple kinases, but 427 were complexed to just a single kinase. Hinge binders typically form 1 to 3 hydrogen bonds to the protein, and while there didn’t seem to be a correlation between the number of hydrogen bonds and potency, more hydrogen bonds did correlate – perhaps counterintuitively – with lower selectivity. To the extent that hydrogen bonds are thought of as enthalpic interactions, this further muddies the argument that enthalpy and entropy can be useful in drug design.

On a more positive note, Sandor Vajda (Boston University) suggested that, according to analyses done in FTMap, perhaps 60-70% of protein-protein interactions may be druggable – as long as we accept that this may require building larger molecules than commonly accepted. And Chris Radoux (Cambridge Crystallographic Data Centre) discussed the computational tool for characterizing hotspots that we previously covered here; a web server for easy search should be available soon.

Library design was also a key topic. Richard Taylor of UCB described his analysis of all FDA-approved drugs, which revealed >350 ring systems. Interestingly though, 72% of drugs discovered since 1983 rely exclusively on ring systems used prior to that date. Clearly there is plenty of untapped chemical real estate.

But getting there won’t necessarily be easy. David Rees stated that 33 fragments recently added to the Astex library required 13 different reaction types. Importantly, many of the fragment to lead successes at Astex have required growing the fragment from the carbon skeleton rather than from more synthetically tractable heteroatoms. Knowing in advance how to do this with every new member of a fragment library should make life much easier in the long run, though it is a serious challenge for chemists.

There is far more to write about, including a great discussion led by Rod Hubbard on how FBLD is integrated effectively into organizations and how it enables difficult targets, but in the interest of space I’ll stop here. If you were at FBLD 2016 (or even if you weren’t) please share your thoughts!

01 February 2016

Fragment-Based Drug Discovery: Lessons and Outlook

In 2006, Wolfgang Jahnke and I co-edited the very first book on fragment-based drug discovery. Half a dozen books have followed, most of which have been reviewed at Practical Fragments (see right-hand column). These are now joined by a new book edited by Wolfgang and me in Wiley’s Methods and Principles in Medicinal Chemistry series.

At 500 pages and 19 chapters, this is the most extensive treatment since the Methods in Enzymology volume five years ago. In the interest of space I can’t write more than a sentence or two about each chapter, but I would like to thank all the contributors. Although I’m undoubtedly biased, I believe this work will set the standard for years to come.

The book is divided into three sections, starting with The Concept of FBDD. Rod Hubbard (Vernalis and University of York) opens with a chapter on the role of FBDD in lead-finding, which provides an introduction, historical overview, and summary of current thinking and future challenges. One particularly interesting section compares the contents of the 2006 book with the state of the art today, highlighting the fact that many of the basic techniques were already in place a decade ago, but the number of success stories has increased dramatically.

Chapter 2, by Glyn Williams and colleagues at Astex, discusses how to choose targets for FBDD, including concepts such as ligandability. Key principles are nicely illustrated with several important targets including the IAPs and HCV-NS3.

The last two chapters in this section focus more on numbers. Chapter 3, by Jean-Louis Reymond and colleagues at the University of Berne, covers the computational enumeration of chemical space, with a special emphasis on the contents and uses of their GDB-17 set of the 166 billion possible molecules with up to 17 non-hydrogen atoms. And chapter 4, by György Ferenczy and György Keseű at the Hungarian National Academy of Sciences, provides an overview of various metrics (such as ligand efficiency and LELP) and how these can be useful for fragment optimization.

The next nine chapters comprise the longest sub-section of the book, Methods and approaches for FBDD. To start screening fragments, you need a library, and designing one is the subject of chapter 5, by Martin Drysdale and colleagues at the Beatson Institute. This chapter also touches on concepts such as molecular complexity and “three-dimensional” fragments.

Screening techniques are best used in combination, and in chapter 6 Ben Davis (Vernalis) and Tony Giannetti (Google[x]) describe the synthesis of results from SPR, NMR, X-ray, ITC, functional screens, and other techniques to overcome challenges in several discovery programs. They emphasize that universal agreement among different methods is not always necessary, but carefully analyzing discrepancies can reveal unexpected problems with the screening conditions, target, or hits.

Differential scanning fluorimetry (DSF) – or thermal shift (TS) – is perhaps the most controversial screening method, and in chapter 7 Chris Abell and colleagues at the University of Cambridge cover this approach in depth. The chapter starts with a thermodynamically detailed yet nonetheless lucid discussion of the theory behind DSF, including the interpretation of negative thermal shifts. The chapter also includes plenty of practical advice and case studies, some of which we’ve covered briefly (for example here and here).

Chapter 8, by Sten Ohlson and Minh-Dao Duong-Thi at Nanyang Technological University, covers three emerging fragment screening technologies: WAC, native MS, and MST. And Chapter 9, by Sandor Vajda (Boston University) and collaborators, does an excellent job of summarizing computational approaches.

As others have noted, some of the biggest challenges are not technical but organizational, and in chapter 10 Michelle Arkin and colleagues at UCSF describe how to make FBDD work in academia. The chapter also includes some interesting polling data, concise but cogent summaries of fragment-finding techniques, and case studies on p97 and caspase-6. And in chapter 11, Jim Wells and colleagues – also at UCSF – describe using Tethering to find allosteric sites in proteins.

One area that has grown dramatically since 2006 is the use of FBDD in complex systems (such as membrane proteins), the subject of a chapter by Miles Congreve and John Christopher at Heptares. Chapter 12 also includes successful case studies, some of which we’ve covered. But finding fragments against these targets is still not easy, as illustrated in the final figure: of 18 fragment hits on 15 targets, almost all have ligand efficiency values > 0.3 kcal/mol per atom, and most of them are relatively potent, with affinities in the mid-micromolar range or better. While everyone wants to find strong binders from the start, such numbers suggest many weak-binding hits are overlooked.

Chapter 13, by Jörg Rademann and colleagues at Freie Universität Berlin, covers protein-templated fragment ligation methods, both reversible and irreversible. The chapter is wide-ranging and includes methods such as dynamic libraries and various types of “Click” chemistries.

The last section of the book, which was mostly absent a decade ago, is entitled Successes from FBDD. This starts with a chapter by Daniel Wyss, Andrew Stamford, and colleagues from Merck on BACE inhibitors. As we’ve noted, fragments have had a major role in most of the BACE inhibitors to enter the clinic, with phase III results from Merck’s verubecestat expected next year.

Epigenetics has also been strongly influenced by fragments, and in chapter 15 Aman Iqbal (Proteorex) and Peter Brown (Structural Genomics Consortium) survey the field, with case studies on several proteins that modulate epigenetic marks. These include BRD4, ATAD2, BAZ2B, SIRT2, and others.

One of the original selling points of fragment-based methods is the ability to go after difficult targets such as protein-protein interactions, and this is the subject of chapter 16, by Feng Wang and Stephen Fesik (Vanderbilt University). In addition to general guidelines, the researchers describe a number of case studies, including RPA, MCL-1, and K-Ras.

Some enzymes can be just as difficult as protein-protein interactions, and in chapter 17 Alexander Breeze (University of Leeds) and former AstraZeneca colleagues describe programs to find inhibitors of LDHA (see here and here). They also discuss how some previously reported inhibitors turned out to be artifacts.

More than two dozen kinase inhibitors have been approved by the US FDA, including the first drug derived from FBDD. In chapter 18, Gordon Saxty (Fidelta) surveys a number of kinase programs, including most of the fragment-derived inhibitors in clinical trials.

And finally, in chapter 19 Simon Rüdisser and colleagues from Novartis present an extensive discussion of renin, with special attention to their campaign, which involved a combination of HTS and fragment-based approaches.

While it may not be possible to judge a book by its cover, the cover of this book does illustrate some of the fruits of the field, with structures of three fragment-derived drugs that have entered the clinic. These are just a small fraction of the 30+ drugs working their way through the pipeline, and of the many more that will spring from the research described and informed by the work presented.

24 June 2015

One Fragment to Rule them All

Recently, I have been riffing on the ontology of FBDD.  FBDD has become so popular that we are now seeing appropriation of the term in many papers that don't really mean it.  So, I came across this paper.  Now, don't be fooled by the title, this is about fragments, the abstract promises me so.  Let me skip the science, which to my eyes is actually quite boring, and get right to the heart of their fragment case.  
How is this paper fragments you ask?  Well, this is not about scaffold hopping or innovative uses of fragments to develop SAR.  This is not about interesting approaches to screening.  It is most certainly not about in silico approaches.  This is most certainly about fragment library design.  We often discuss here the sizes of fragment libraries and what they should look like.  One important concept we often tackle here is how big should the libraries be and what size should fragments be.  More importantly we often discuss how much of chemical space a fragment library should cover.  This paper takes an anti-Reymond approach to address that question. 
The Reymond approach tries to determine how big chemical space is, what it looks like, and what portion of it is available.  The Anti-Reymond approach identifies what is available and validates its inclusion in a fragment library.  Here is the last sentence of this paper:
"These findings...verify the value of the benzamide fragment in drug design."
Now, I was worried that benzamidine was not a valuable fragment.  This paper has removed all doubt in my mind.  Now that is settled, we can go on an validate the other 165, 999, 999,999 other possible fragments. 

30 June 2011

How effectively can fragments sample chemical space?

One of the key advantages of fragment-based drug discovery is that, since there are fewer fragments than lead-sized or drug-sized molecules, it is possible to sample chemical space far more efficiently with fragments than with larger molecules. At least, that’s the theory, but does is it hold true in the real world?

To put it another way, do fragments sample all of the space in which drugs are found? And what kinds of fragments are best for this sampling? In the most recent issue of J. Med. Chem., Stephen Roughley and Rod Hubbard of Vernalis address such questions.

The system they investigate, heat shock protein 90 (Hsp90), is an ideal model system: it is both a popular anti-cancer target as well as structurally tractable, and is thus arguably the most heavily explored single target in terms of fragment-based lead discovery. At least 8 antagonists have entered the clinic, of which at least 2 have come from fragments (see the posts on AT13387, NVP-BEP800/VER-82576, and posts on Evotec compounds discovered by fragment growing or linking.)

Vernalis has had a long-running fragment-based program targeting Hsp90, which has resulted in numerous fragments whose binding modes have been determined by X-ray crystallography. Roughley and Hubbard analyzed these fragments and compared them to published inhibitors. Just 5 distinct fragments can be mapped onto all of the clinical compounds: a handful of fragments effectively samples relevant chemical space. As the authors put it:

For Hsp90 at least, the fragments do cover an appropriate chemical space; what is then important is the imagination of the chemist in evolving the fragments into potent inhibitors.

The second point – about the imagination of the chemist – is critical. Mapping fragments onto elaborated molecules is easier to do retrospectively than prospectively; a cynic could argue that methane is a fragment of just about any drug out there. However, Roughly and Hubbard also point out that, particularly in cases such as this where there are many co-crystal structures, fragments can help identify bioisosteres, including cryptic ones that would not be obvious purely from studying functional SAR.

The paper also addresses the issue of optimal library design, in particular the dilemma of size. Although all five representative fragments were found in an initial set of just 719 fragments, subtle changes can dramatically change the binding mode, an issue we’ve touched on previously. It may not be practical to have multiple similar fragments present in a primary screening library, but testing close analogs after identifying initial fragment hits is likely to be worthwhile.

Finally, one of the concerns about fragment-based approaches is that, if everyone is buying the same set of fragments from the same suppliers and screening them against the same targets, they will end up in the same place – and stumbling over each others’ intellectual property. Reassuringly, this turns out not to be the case:

Even though the various companies discovered rather similar compounds from a fragment screen, exploiting similar binding motifs, there were no exact matches. [Also], the subsequent evolution of the fragments sometimes took very different paths and produced mostly very different chemical leads and candidates.

If this holds true for such heavily mined targets as Hsp90 (and kinases, as discussed previously) it should be even more true for newer classes of targets.

There is a wealth of information in this paper, and it is worth perusing, especially if you find yourself longing for some science over the long holiday weekend.

14 October 2010

FBLD 2010

The last major fragment event of this year is over, but it ends on a high note: FBLD 2010 has remained true to its predecessors in bringing together a great group of fragment enthusiasts in a Gordon-Conference-like environment. With 30 talks and even more posters I won’t attempt to be comprehensive or even representative, but will instead just pick out a few themes. Those of you who were there, please chime in with your own observations.

One of the themes was the shape of chemical space, and what makes a good binder. Jean-Louis Reymond, who has been systematically enumerating all stable molecules containing carbon, nitrogen, oxygen, and a few other atoms, has already published up to 13 heavy atoms but is now expanding his analysis to molecules containing up to 17. The issue of whether more attention should be given to three-dimensional fragments was discussed, with Ken Brameld reporting that fragments in crystal structures at Roche and the protein data bank contain fewer “flat” compounds than does the ZINC database of commercial molecules. However, analyzing 150,000 molecules with MW < 300 that had been screened in 40-100 high-throughput screens at Roche did not show any shape differences between the 50,000 molecules that showed up in at least one screen and those that didn’t. Interestingly, this ratio also came up in a talk by Tony Giannetti of Genentech, who said that across 13 screens 36% of their fragments hit at least one protein, while the rest didn’t hit any. Vernalis has found similar results; is there any way to enrich for the productive binders?

While FBLD 2009 had a strong computational theme, a major thrust of this conference was using biophysics to detect and confirm fragment binding. Tony discussed best-practices in SPR, and noted that since small molecules are “brighter” in NMR assays than proteins it is possible to find even very tiny fragments, including a 6 heavy-atom compound with a Kd of 600 micromolar. Tony also described the use of SPR for weeding out badly behaved compounds. Spookily, he noted that promiscuity is a function of compound, protein, and buffer, so it is not possible to weed out bad actors in a library before screening: one compound that was promiscuous against 8 targets bound legitimately and gave a crystal structure with a ninth. Adam Renslo of the University of California San Francisco described how easy it is to be misled by such phenomena. Glyn Williams described how Astex uses biophysical techniques to detect problem compounds, and noted that oxidizers can be particularly insidious – a trend that will likely continue as people explore novel heterocycles.

Glyn also presented a fascinating if slightly depressing discussion of ligand efficiency. As many have found, it can be challenging to maintain ligand efficiency during the course of fragment optimization. Yet even this goal is too modest. A fragment pays about 4.2 kcal/mol in binding energy when it binds to a protein due to loss of rotational and translational entropy; since this enropy cost is only paid once, atoms added to this molecule do not have this liability . Thus, merely maintaining ligand efficiency means that the atoms being added are binding less efficiently. This point was also emphasized by Colin Groom of the Cambridge Crystallographic Data Centre.

Membrane proteins are increasingly being targeted by fragment-based methods, as recently discussed on this site, and both Gregg Siegal of ZoBio and Rebecca Rich of the University of Utah presented progress against GPCRs.

There was general agreement that many approaches can find fragments, and that using several orthogonal methods is a good way to separate the true binders from the chaff, but a continuing challenge is what to do next. The last two sessions were devoted to chemical follow-up strategies and success stories. Some of these have been at least partially covered on Practical Fragments (for example here, here, and here) but there were a number of unpublished examples too – we’ll try to discuss these individually as they emerge.

If you missed this or the previous two conferences you’ll have another chance in 2012, when the meeting will be held in my fair city of San Francisco. And if you can’t wait that long, there are at least two fragment conferences scheduled for next year – details to come shortly.