Published today in Nature Biotechnology, AdaptiveFlow makes ultra-large virtual drug screenings more accessible, scalable and efficient, with a 1,000-fold reduction in computational costs over existing methods.
Researchers today announce AdaptiveFlow, an AI-informed platform that can virtually screen billions of drug-like molecules with a 1,000-fold reduction in computational costs over existing methods. Developed and validated by scientists from St. Jude Children’s Research Hospital, University of Pavia, Dana Farber Cancer Institute and Harvard Medical School, AdaptiveFlow allows prohibitively expensive ultra-large virtual drug screens to be done routinely. The open-source platform was published today in Nature Biotechnology.
The platform’s framework demonstrated linear scaling up to 5.6 million virtual central processing units (CPUs) — a new benchmark for cloud-based drug discovery — allowing billions of molecules to be screened without loss of efficiency. As proof of concept, the team identified potent inhibitors for existing and emerging cancer targets for which few inhibitors are known.
“AdaptiveFlow is the next generation in automated drug discovery platforms for routine ultra-large virtual screenings,” said co-corresponding author Christoph Gorgulla, PhD, Center of Excellence for Data-Driven Discovery, St. Jude Department of Structural Biology. “With this platform, we are able to screen 69 billion molecules, representing the largest ready-to-dock library in the world.”
Multidimensional grid allows for multibillion-molecule screening
The recent expansion of ultra-large molecule screening libraries provided a call to action for the drug discovery field to unlock their potential. While pioneering efforts saw success in billion-compound screens, AdaptiveFlow is the first of a new generation focused on efficiency, affordability and access. This is rooted in its economical approach to computation.
“A lot of software loses communication efficiency with increasing CPU count,” Gorgulla explained, “but with AdaptiveFlow, the scaling behavior is perfectly linear, even with millions of CPUs. That is special.”
At the core of AdaptiveFlow is an 18-dimensional grid in which each dimension represents a specific molecular property, such as molecular weight. This framework allows researchers to prioritize chemically diverse molecules and guides the rational selection of promising library subsets for deeper screening. A machine-learning classification model trained on these prescreening results then identifies the most promising molecules. These are then virtually screened against the protein target with over 1,500 supported docking protocols, which are subsequently ranked according to their predicted binding affinity.
“The initial hits are often of considerably better quality than traditional methods, which can save researchers much time and effort during the optimization phases of drug discovery,” Gorgulla said.
AdaptiveFlow identifies strong inhibitors for elusive target
To demonstrate its potential, the team used AdaptiveFlow to find inhibitors for a well-known anticancer target, poly(ADP-ribose) polymerase 1 (PARP1), and an emerging target, ferroptosis suppressor protein 1 (FSP1), which is involved in cell death and has recently been linked to cancer cell survival. In both cases, the platform identified inhibitors exhibiting binding strength comparable to pharmaceutically relevant standards.
“There are existing approved drugs on the market for PARP1, which made it a great benchmark for us, but FSP1 is a more challenging target because its binding site has an additional co-factor present,” Gorgulla explained. “We wanted to see if AdaptiveFlow also works in this complex setting and were excited to see it does.”
The successful identification of FSP1 inhibitors underscores AdaptiveFlow’s potential as a powerful tool for modern drug discovery. By democratizing access to ultra-large molecule libraries, faster and more impactful therapeutic development can follow.
“We have step-by-step guidance and tutorials available, and the entire platform is open source, so that everyone can use it for free,” Gorgulla said. “We wanted to make ultra-large virtual screening as accessible as possible so it can translate into better molecules for more challenging target proteins that can really help improve clinical success rates.”
AdaptiveFlow is available for download at https://adaptive-flow.ai/ and https://github.com/QuantumAI4Bio
Authors and funding
The study’s co-corresponding authors are Andrea Mattevi, University of Pavia; and Haribabu Arthanari, Dana Farber Cancer Institute and Harvard Medical School. The study’s co-first authors are Domiziana Cecchini, University of Pavia; AkshatKumar Nigam, Klyne and Stanford University; Ming Tang, The University of Queensland, Queensland University of Technology and St. Jude; and Joana Reis, Dana-Farber Cancer Institute. The study’s other authors are Matt Koop, Amazon Web Services; Andrea Gottinger and Callum Robert Nicoll, University of Pavia; Abhilash Jayaraj, Dana Farber Cancer Institute and Harvard Medical School; Süleyman Selim Çınaroğlu, University of Oxford and Ege University; Ricarda Törner, Dana Farber Cancer Institute and University of Zurich; Yehor Malets, National Academy of Science of Ukraine; Minko Gehev, Google; Krishna Padmanabha Das, Harvard Medical School and St. Jude; Hyuk-Soo Seo, Sirano Dhe-Paganon, Eun-Bee Choi, Geoffrey Shapiro, Huel Cox 3rd, Luke Sebastian, Chelsea Braithwaite and Puspalata Bashyal, Dana Farber Cancer Institute; Christopher Secker, Zuse Institute Berlin (ZIB) and Max Delbrück Center for Molecular Medicine; Mohammad Haddadnia, University of Toronto; Alexander Hasson, University of Oxford; Minkai Li, Harvard College; Abhishek Kumar, Indian Institute of Technology Guwahati and St. Jude; Roni Levin‑Konigsberg, Klyne; Dmytro Radchenko, Enamine Ltd; Aditya Kumar, Technical University Berlin; Pierre-Yves Aquilanti, Amazon Web Services and NVIDIA; Henry Gabb, Intel; Amr Alhossary, Wesleyan University and Pulse for Integrated Solutions GmbH; Gerhard Wagner, Harvard Medical School; Alán Aspuru-Guzik, University of Toronto, Vector Institute for Artificial Intelligence and Canadian Institute for Advanced Research; Yurii Moroz, Enamine Ltd, Chemspace LLC and Taras Shevchenko National University of Kyiv; Konstantin Fackeldey, Zuse Institute Berlin (ZIB) and Technical University Berlin; and Yao Wang, Kelly Churion, Jongwan Kim, Nidhin Thomas, Yong Li, Lei Yang, Charalampos Kalodimos and John Schuetz, St. Jude.
The study was supported by the National Institutes of Health (GM136859), the Bio-X Stanford Interdisciplinary Graduate Fellowship (SGIF), the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) under Germany´s Excellence Strategy – The Berlin Mathematics Research Center MATH+ (EXC-2046/1, project ID: 390685689), J. Goldberg, the European Research Council (ERC Advanced Grant, MetaQ no. 101094471), the Associazione Italiana per la Ricerca sul Cancro (AIRC) Investigator Grant (28754) and the American Lebanese Syrian Associated Charities (ALSAC), the fundraising and awareness organization of St. Jude.
St. Jude Children's Research Hospital
St. Jude Children’s Research Hospital is leading the way the world understands, treats, and cures childhood catastrophic diseases. From cancer to life-threatening blood disorders, neurological conditions, and infectious diseases, St. Jude is dedicated to advancing cures and means of prevention through groundbreaking research and compassionate care. Through global collaborations and innovative science, St. Jude is working to ensure that every child, everywhere, has the best chance at a healthy future. To learn more, visit stjude.org, read St. Jude Progress, a digital magazine, and follow St. Jude on social media at @stjuderesearch.