AI speeds up drug discovery by predicting what used to be found by trial and error: which protein to target, what that protein looks like, which molecules might bind it, which will be toxic, and which patients fit a trial. It shortens the early search. As of September 2026, no drug discovered mainly by AI has been approved.

This guide goes through the pipeline stage by stage: the machine learning methods used for target identification, protein structure prediction with AlphaFold, virtual screening and generative chemistry, ADMET and toxicity prediction, and clinical trials. It then sets out what AI-discovered drugs have achieved in people, how the FDA and the EMA treat AI, and where the approach falls short. Physics-based simulation and quantum chemistry, the other branch of computational drug design, have a post of their own. This article describes research and regulation; it is health information, not medical advice.

Why drug discovery is slow, and where AI fits

A small-molecule drug starts as a hypothesis: acting on a particular protein, the target, will change the course of a disease. Chemists then look for molecules that bind the target (hits), improve the best ones into leads, and check them for absorption, metabolism and toxicity before a single candidate is given to people. Clinical trials follow in three phases, and that is where most candidates fail. The FDA's patient guide to clinical research puts the share of drugs that move on at about 70% after phase 1 and about 33% after phase 2, and a phase 3 study alone lasts one to four years.

Machine learning helps in two ways. It shrinks the search, by scoring millions or billions of options so that only the most promising few are made and tested. And it tries to predict failure earlier, before money is spent on synthesis, animal studies or patients. The FDA's discussion paper on AI and machine learning in drug development, first published in May 2023 and revised in February 2025, describes uses at every stage:

StageThe questionWhat machine learning doesExample below
Target identificationWhich protein drives the disease?Mines omics data and the literature to rank candidate targetsTNIK in lung fibrosis
Structure predictionWhat does the target look like?Predicts a 3D structure from the amino acid sequenceAlphaFold 2 and AlphaFold 3
Hit finding and designWhich molecules bind it?Scores huge libraries, generates new moleculesDeep Docking, GENTRL, Chemistry42
ADMET and toxicityWill it reach the target without harm?Predicts absorption, metabolism and toxicity from structureTox21 challenge, DeepTox
Clinical trialsWho, where and how many?Matches patients, ranks sites, predicts outcomes on placeboTrialGPT, PROCOVA

Target identification: choosing what to hit

Everything downstream depends on the target. A wrong one wastes years, because a molecule can bind it perfectly and still do nothing for the disease. At this stage AI works mostly as a reading and ranking machine: models combine genomic, transcriptomic and proteomic data from healthy and diseased tissue with the published literature, and rank proteins by how strongly the evidence ties them to the disease. The FDA's discussion paper describes exactly this use, and adds the caution that matters: the target's role in the disease still has to be validated in later studies.

A well-documented example is TNIK in idiopathic pulmonary fibrosis (IPF), a progressive scarring of the lungs. In a Nature Biotechnology paper published on 8 March 2024, Insilico Medicine described running its PandaOmics platform, which combines several AI engines including generative pretrained transformers, over fibrosis datasets; TNIK came out as the top-ranked candidate. The team then designed inhibitors with its Chemistry42 platform (next sections), showed anti-fibrotic activity in several mouse and rat models, and ran phase 1 safety trials, one of them with 78 healthy participants. The authors report that the work took roughly 18 months from target discovery to nominating a preclinical candidate.

Protein structure prediction: AlphaFold and AlphaFold 3

Structure-based drug design needs the target's three-dimensional shape, and in particular the pocket where a drug could bind. Experiments determine it by X-ray crystallography, NMR or cryo-electron microscopy, but slowly. When the AlphaFold 2 paper appeared in Nature on 15 July 2021, experiments had determined the structures of around 100,000 unique proteins, a small fraction of the billions of known sequences, and one structure could take months to years of work. AlphaFold 2 was validated in CASP14, a blind assessment, where its predictions were competitive with experimental structures in a majority of cases.

In October 2024, half of the Nobel Prize in Chemistry went to Demis Hassabis and John Jumper of Google DeepMind for protein structure prediction. The Nobel committee noted that AlphaFold 2 had been used to predict the structure of virtually all of the 200 million proteins researchers have identified, by more than two million people in 190 countries. How AlphaFold fits among other AI-driven results in science is covered in AI-enabled scientific discoveries.

For drug discovery, AlphaFold 3, published in Nature on 8 May 2024, matters more, because it predicts complexes: proteins together with small molecules, nucleic acids, ions and modified residues. On the PoseBusters benchmark, 428 protein-ligand structures released to the Protein Data Bank in 2021 or later, it placed ligands far more accurately than classical docking tools such as AutoDock Vina, without being given the protein's experimental structure. Its authors are plain about the limits. Like earlier predictors, it produces static structures of the kind stored in the PDB, not how molecules move in solution, and as a generative model it can hallucinate plausible-looking structure in disordered regions.

A straight chain of grey beads passes through a gear and comes out folded into a compact ribbon with a pocket, where a small orange ring-shaped molecule sits.
Fig. 1 A predicted structure gives chemists a pocket to design against, but it is one still frame of a moving protein, so binding still has to be measured.

Virtual screening and generative chemistry

Once there is a target and a pocket, the question is which molecule to make. Chemical libraries now run to billions of molecules, many of them make-on-demand compounds that have not been synthesized yet, far more than any lab could test. Two families of methods address that.

Virtual screening: score everything, make a few

Virtual screening ranks a library by how well each molecule is predicted to bind. In a 2019 Nature study, researchers docked 170 million make-on-demand compounds against two targets, then synthesized and tested 44 and 549 of the top-ranked molecules. The screen found inhibitors with no known precedent, and crystal structures confirmed the docking predictions for the enzyme target.

Docking is a physics-style calculation, and at billions of molecules it becomes the bottleneck, so machine learning is used to decide what is worth docking. Deep Docking (ACS Central Science, 2020) trains a deep learning model on the docking scores of a sample of the library, uses it to discard unpromising molecules, and repeats. With it, the authors covered 1.36 billion molecules from the ZINC15 library against 12 target proteins, with up to 100-fold data reduction.

Hundreds of small molecule symbols pass through a gear into a narrowing funnel. Six drop into test tubes in a rack, and one orange molecule leaves the last tube with a check mark.
Fig. 2 The model decides what is worth making. Only the few compounds that are synthesized and measured count as hits.

Generative chemistry: design molecules that do not exist yet

Generative models work the other way round. Instead of ranking a list, they propose new structures optimized for several goals at once, such as predicted activity, novelty and ease of synthesis. In a 2019 Nature Biotechnology paper, Insilico's GENTRL model (generative tensorial reinforcement learning) produced inhibitors of the kinase DDR1 in 21 days. Four compounds were active in biochemical assays, two were validated in cell-based assays, and one lead showed favourable pharmacokinetics in mice. The TNIK inhibitor that became rentosertib also came from generative chemistry: the Chemistry42 structure-based design workflow working from crystal structures of TNIK's kinase domain and its ATP-binding pocket.

Look at the ratio in each of these studies: millions or billions of candidates in, dozens or hundreds made, a handful confirmed. The model's job is to raise the lab's hit rate, not to replace the lab.

ADMET and toxicity prediction

A molecule that binds its target still has to reach it and do no harm. ADMET covers absorption, distribution, metabolism, excretion and toxicity. Machine learning models predict these properties from chemical structure, so that compounds likely to be poorly absorbed, quickly cleared or toxic can be dropped before anyone makes them.

The benchmark that showed deep learning could do this was the Tox21 Data Challenge: 12,000 environmental chemicals and drugs, measured for 12 different toxic effects in purpose-built assays. DeepTox, a pipeline from Johannes Kepler University Linz, had the highest performance of all computational methods and won the grand challenge. The authors credit part of that to multi-task learning: one network learns all the toxic effects at once, and in doing so learns highly informative chemical features.

Regulators have started to encourage this kind of evidence. On 10 April 2025 the FDA announced a plan to reduce, refine or potentially replace animal testing in the development of monoclonal antibodies and other drugs, using approaches that include AI-based computational models of toxicity alongside cell lines and organoids. Such data was encouraged in investigational new drug (IND) applications from that day.

Note

Prediction reduces surprises; it does not remove them. In rentosertib's phase 2a trial, described below, the most common events that led patients to stop treatment were related to liver toxicity or diarrhea.

AI in clinical trials: design, site selection and recruitment

Trials take years, and AI does not make biology respond faster. What it can do is make a trial better aimed and quicker to fill. The FDA's discussion paper lists the main uses:

  • Recruitment: mining trial registries, electronic health records, the literature and other sources to match individuals to trials, while making sure the populations likely to use the drug are adequately represented.
  • Selection and stratification: predicting an individual participant's outcome from baseline data such as demographics, lab results, imaging and genomics, so a trial can enrol the patients most likely to show an effect.
  • Site selection: evaluating how sites performed in other trials, to pick those with the best potential and flag those likely to run behind schedule.
  • Digital twins: an emerging method in which a model builds a computational representation of an individual that tracks their molecular and physiological status over time.

Recruitment suits language models, because eligibility criteria and patient notes are both text. TrialGPT, from the US National Library of Medicine and collaborators, published in Nature Communications on 18 November 2024, retrieves candidate trials for a patient, checks each eligibility criterion against the patient's notes and ranks the trials. Evaluated on 183 synthetic patients, it reached 87.3% accuracy on criterion-level eligibility, close to expert performance, and in a user study it cut screening time by 42.6%.

On trial design, the European Medicines Agency took a concrete step in September 2022, when its human medicines committee (CHMP) qualified PROCOVA, prognostic covariate adjustment, for phase 2 and 3 trials with continuous outcomes. A model trained on historical data predicts each participant's outcome on placebo, and that prediction enters the analysis as a covariate, which can increase statistical power or, with care, allow a smaller sample. The opinion also warns that complex models, machine learning included, risk overfitting and should be checked on data independent of their training data.

What AI-discovered drugs have achieved by 2026

Pipeline claims only count once a drug works in people. Here is where the evidence stands as of September 2026.

Rentosertib: phase 2a results, then phase 3

Rentosertib (formerly ISM001-055) is the TNIK inhibitor from the sections above: both its target and the molecule came from generative AI. Its phase 2a trial, published in Nature Medicine on 3 June 2025, randomized 71 adults with IPF at 21 sites in China to one of three doses or placebo for 12 weeks. The primary endpoint was safety: the share of patients with at least one treatment-emergent adverse event was similar across arms, 70.6% on placebo and 72.2% to 83.3% on rentosertib. Lung function, a secondary endpoint, moved the right way at the highest dose: forced vital capacity rose by a mean of 98.4 mL on 60 mg once daily and fell by 20.3 mL on placebo. The authors name the limits themselves: small arms, participants who all lived in China, and short follow-up.

A phase 3 trial (NCT07687459) started on 9 September 2026, according to its ClinicalTrials.gov record. It runs 52 weeks against placebo, with an estimated 320 participants and primary completion estimated for October 2029.

REC-994: a phase 2 readout that did not hold

Recursion reported phase 2 results for REC-994 in cerebral cavernous malformation on 3 September 2024. REC-994 is a small molecule whose potential in the disease was first shown with an early version of the technology behind Recursion's platform. The 12-month SYCAMORE trial met its primary endpoint of safety and tolerability and showed a trend toward smaller lesions on MRI at the highest dose, but no improvement yet in patient- or physician-reported outcomes. With its first-quarter 2025 results on 5 May 2025, Recursion said the totality of the data supported discontinuing the study: long-term extension results showed no promising trends, and the 400 mg arm was indistinguishable from natural history.

Success rates, and why there is no approval yet

A 2024 analysis in Drug Discovery Today, by Boston Consulting Group authors, looked at the clinical pipelines of AI-native biotech companies. AI-discovered molecules had an 80% to 90% success rate in phase 1, substantially higher than historic industry averages, and about 40% in phase 2, on a limited sample, comparable to historic averages. The authors read this as a sign that AI is highly capable of producing molecules with drug-like properties. Whether a drug helps patients depends on more than that, starting with the choice of target, and only trials show it.

No drug discovered mainly by AI has been approved yet. A review published online on 27 June 2026 in Critical Reviews in Oncology/Hematology states that no fully AI-discovered and AI-designed drug has received marketing approval, even though the FDA has approved numerous AI-enabled medical devices and software tools.

Important

"Months instead of years" claims refer to discovery, the stage before trials. Rentosertib shows the real scale: about 18 months from target to preclinical candidate, then phase 1, a 12-week phase 2a, and a phase 3 whose primary completion is estimated for October 2029.

How regulators treat AI in drug development

On 6 January 2025 the FDA published draft guidance on using AI to support regulatory decisions for drugs and biological products. The draft proposes a risk-based credibility assessment in seven steps:

  1. Define the question of interest the AI model will address.
  2. Define the model's context of use.
  3. Assess model risk, which combines model influence (how much the AI output weighs against other evidence) and decision consequence (how bad an incorrect decision would be).
  4. Develop a plan to establish the credibility of the model's output for that context of use.
  5. Execute the plan.
  6. Document the results and discuss any deviations from the plan.
  7. Decide whether the model is adequate for the context of use.

Two details matter for this post. First, the guidance explicitly excludes AI used in drug discovery, and AI used for operational efficiency that does not affect patient safety, drug quality or the reliability of study results. It covers AI whose output supports decisions on safety, effectiveness or quality. Second, it is still a draft: as of September 2026, the FDA's guidance page lists it as draft guidance, not for implementation. The FDA's drug development AI page says the draft drew on CDER's experience with more than 500 submissions with AI components from 2016 to 2023.

In January 2026 the FDA and the European Medicines Agency published 10 guiding principles of good AI practice in drug development, among them a risk-based approach, a clear context of use, data governance and documentation, and life cycle management. The EMA's reflection paper on AI in the medicinal product lifecycle, adopted in September 2024, takes a similar line on discovery: AI used there can be of low regulatory impact when poor performance only affects the developer, but once its results become part of the evidence submitted for review, the principles for non-clinical development apply.

The limits: data quality, lab validation and hype

Data quality and leakage

A model learns what its data records, including its gaps and duplicates. In binding affinity prediction, a study in Nature Machine Intelligence, published on 21 October 2025, found train-test data leakage between the widely used PDBbind database and the CASF benchmark sets that had severely inflated the reported performance of deep learning models. When the authors retrained top models on a cleaned split, their benchmark performance dropped substantially.

Two separate stacks of cards, training data and test data. One orange card in the test stack is joined by a dashed line to a matching card in the training stack, beside a gauge pushed to its maximum.
Fig. 3 When test examples repeat the training data, a model is graded on answers it has already seen, and its score says little about new targets.

Regulators raise the population side of the same problem. The FDA's discussion paper stresses adequate representation of the populations likely to use the drug, and the EMA asks sponsors to review models and datasets for bias, including discrimination against non-majority genotypes.

Validation in the lab

A prediction is a hypothesis. The PoseBusters study (Chemical Science, 2024) checked five deep learning docking methods for physical plausibility, such as bond lengths, stereochemistry and clashes with the protein, and found that none yet outperformed classical docking tools on plausibility or on examples unlike their training data. AlphaFold 3 later reported beating classical docking on the PoseBusters set, but a predicted pose is still a static shape that binding assays, crystal structures and, eventually, animal studies and trials have to confirm. Every result in this article passed through that filter: GENTRL's 21 days produced four active compounds, not a medicine.

Hype against results

A fair summary of AI in drug discovery in 2026: faster searches, a strong phase 1 record, a phase 2 record like everyone else's, no approvals yet, and at least one candidate stopped after phase 2. When you read a claim, check five things:

  1. Which stage it refers to. Faster discovery is not clinical success.
  2. Whether the prediction was tested. A docking score or a predicted structure is not a binding measurement.
  3. Whether the benchmark was clean. Look for test data kept apart from training data, by date or by structural similarity.
  4. What the trial measured. Phase 2a primary endpoints are often safety; efficacy trends in small arms need a larger trial.
  5. Who reports it. A company announcement and a peer-reviewed paper are both sources, with different levels of review.

What the pattern means outside pharma

The method generalizes. A model scores a large space of options, people test the few it ranks highest, and test data is kept out of training. The same discipline applies to AI in any organization: the data behind a model decides what it can learn, and an evaluation on real cases decides whether a change ships. If you are building that foundation, our data platform service covers pipelines with tests and lineage and delivers features and datasets for machine learning, and our AI and automation service runs an evaluation set built from your real cases before any model, prompt or source change goes live. Neither is a drug discovery service. For where language models in particular fall short, see AI limitations in understanding.