
Searching beyond an existing compound shelf
Drug discovery often begins by looking for a molecule that affects a protein involved in disease. Existing libraries contain many compounds, but they cover only a small part of the possible chemical structures. A generative model can suggest structures that are not already sitting in a screening collection. The hard part is finding suggestions that chemists can make and that work in a physical experiment. [1]
The researchers built TamGen, a chemical language model that reads molecules as strings of chemical symbols. It was pretrained on 10 million compounds, then given the three-dimensional shape of a target protein's binding pocket. A second part of the system could refine a proposed molecule while retaining useful features from an earlier candidate. [1]
Thousands of suggestions became hundreds of tests
The team focused on ClpP, a protein that tuberculosis bacteria need for protein control. TamGen first generated 2,612 unique molecules. Computer screening reduced these to four starting structures. After another round of generation around those structures and three earlier weak inhibitors, the model produced 8,635 more molecules. The researchers selected 296 for the next stage. [1]
The first physical tests did not begin by synthesizing every generated molecule. The team searched a commercial library for similar compounds, found 159 close matches and tested them against purified ClpP protein. Five inhibited the protein. The researchers then synthesized three new generated molecules related to those hits and later tested eight molecules produced directly by TamGen. [1]
Fourteen early hits, with an important limit
14 active compounds
Five commercial analogues, three synthesized derivatives and six of eight directly generated molecules inhibited the selected protein in laboratory assays.
1.9 micromolar
The concentration at which the strongest tested compound cut the measured protein activity by half. This is a biochemical result, not a treatment dose.
The result matters because the work did not stop at a docking score on a computer. Chemists could obtain or synthesize the selected structures, and physical assays measured an effect on the intended protein. That makes these compounds starting points for further drug research. It does not show that they kill tuberculosis bacteria, work in animals or help people. [1]
What must happen before this resembles a medicine
A protein-level hit can fail for many reasons. The compound may be toxic, break down too quickly, fail to reach bacteria inside the body or affect other proteins. The study did not report tests in tuberculosis-infected cells, animals or people. It also used one protein target, so the reported hit rate cannot be assumed for other diseases. [1]
The model helped propose and prioritize molecules. Human researchers chose the pipeline, applied additional computer filters, found purchasable analogues, synthesized compounds and ran the experiments. The progress comes from that combined process, not from an autonomous AI discovering a drug. [1]
The authors released code, model weights and supporting data. That makes the computational work easier to inspect and repeat, but an independent laboratory would still need to reproduce the biological results and test the compounds in more realistic systems. [1]
Sources & context
The PMC article and publisher DOI are two routes to the same peer-reviewed study. The linked repository contains supporting code, weights and data, not an independent validation.
TamGen: drug design with target-aware molecule generation through a chemical language model
Wu and colleagues · Nature Communications · October 29, 2024