08 September 2026

Computational prediction of cryptic pockets

Last week we highlighted how NMR-based fragment screening was able to identify cryptic pockets on a small protein. Experimental methods are (still) the gold standard, but they take considerable effort and resources. In a new open-access J. Am. Chem. Soc. paper, Neha Vithani, David LeBard, and collaborators at OpenEye, Bayer, Genentech, and Khartis provide a computational approach.
 
The researchers defined six categories of cryptic pockets: 1) side-chain motion, 2) loop motion, 3) secondary-structure motion, 4) secondary-structure change, 5) interdomain motion, and 6) pre-existing deep pockets with no access to the surface of the protein. They collected a set of 21 proteins with 26 cryptic pockets overall (some proteins contain more than one) whose apo (unbound) and holo (ligand-bound) structures were available in the Protein Data Bank. This “Encore Dataset” contains all six types of cryptic pockets, though some are represented by only two or three examples.
 
To computationally identify cryptic pockets, the researchers used weighted ensemble molecular dynamics (WEMD) simulations on the apo structures. The details are quite technical but involve focusing on “normal modes” of protein movement. To find pockets, xenon is included in the simulation; these large atoms will fill pockets that open up during the course of the simulation, so detecting bound xenon reveals pockets.
 
WEMD successfully identified 24 of the 26 cryptic pockets in the Encore Dataset. Interestingly, both failures were for the same protein, hepatitis C polymerase, one a preexisting deep pocket and the other a secondary-structure change from order to disorder. (This latter category is exemplified in one of the cryptic pockets we discussed last week, in which crystallography revealed that a helix became disordered upon ligand binding.) The researchers hypothesize that these types of order-to-disorder transitions may be particularly difficult to detect computationally because the models are designed to identify secondary structures, thus overstabilizing them.
 
So WEMD is able to correctly predict most cryptic pockets in the Encore Dataset, but does it also predict cryptic pockets that don’t exist? Paul Samulsen famously quipped that economists had “predicted nine out of the last five recessions,” which actually looks pretty good compared to WEMD, which predicted more than 30 potential cryptic pockets on the protein biotin carboxylase. For comparison, the software program fpocket, which we wrote about last year, predicts 13 for this protein.
 
To rank which pockets actually exist, the researchers turned to another computational technique called Target X, which was described in this Research Square article earlier this year. This proved reasonably successful at ranking experimentally observed pockets near the top, with the top two predicted biotin carboxylase pockets corresponding to known pockets.
 
This is an interesting paper, but I am left with more questions. First, while the Encore Dataset is nice, why didn’t the researchers (also) use CryptoSite, a previously published set of 93 proteins with cryptic sites, which we discussed in 2024? More importantly, how can drug hunters decide whether to target a specific cryptic pocket? For instance, does WEMD and Target X allow one to assess whether cryptic pockets can support high-affinity ligands? These would be interesting questions for a follow-up paper.