
Cancer cells do not stand still after a drug arrives
Many models learn from a single snapshot of a cell. This study asked a different question: can a model learn the sequence of protein changes that follows drug exposure? Proteins matter here because most medicines act on proteins, while changes elsewhere in the cell can weaken or redirect a drug's effect. [2]
The researchers treated 18 breast-cancer cell lines with 63 drugs at several time points. They collected 16,311 perturbation-proteomics samples, producing more than 38 million protein-abundance measurements. These are laboratory measurements from cell systems, not measurements from 38 million people or tumours. [2]
The model follows change through time
ProteinTalks uses a neural ordinary differential equation, a method designed to represent change over time. It combines the starting protein profile, information about a drug and measured responses at later time points. The model was then used to predict drug activity and combinations, and to identify proteins associated with resistance. [1] [2]
AI organized a large set of time-based measurements and ranked possibilities. Researchers still chose the experiments, produced the protein data, tested drug combinations and interpreted the results. [2]
Useful predictions depended on what the model had seen
The team tested 98 additional anti-cancer compounds in four breast-cancer cell lines. Across all of those compounds, the preprint reports an accuracy of 0.619 and an AUROC of 0.671. After compounds with mechanisms absent from the training data were excluded, the reported values rose to 0.844 and 0.840. That gap is important: the model handled familiar kinds of drug action better than unfamiliar ones. [2]
The researchers also tested four model-prioritized drug combinations in cell assays and used gene interference experiments to examine proteins linked to drug response. These experiments move the work beyond a purely computational benchmark, but they remain laboratory tests. [2]
The final paper reaches more realistic samples
The peer-reviewed Nature paper reports evaluation beyond cell lines, including patient-derived organoids and clinical biopsy samples. Its abstract says the model generally performed better than the selected comparison methods under the study's protocols. This is evidence that the learned patterns can transfer to more complex material, not proof that the model should choose a person's treatment. [1]
What to keep in view
The training data centre on breast-cancer cell lines and a bounded set of drug mechanisms. The preprint also notes that its high-throughput method measured fewer proteins per sample than some deeper proteomics studies and did not model post-translational modifications. Those choices made the dataset much larger, but left parts of cell biology outside the model. [2]
Sources & context
The bioRxiv manuscript is the authors' public preprint of the same ProteinTalks study and supplies the accessible methods and results. The Nature version of record has a changed title, expanded author list and added validation, so final-version claims above are limited to material visible on the publisher page.
An operational perturbation proteomics-based virtual cell model
Sun and colleagues · Nature · September 9, 2026
A perturbation proteomics-based foundation model for virtual cell construction
Sun and colleagues · bioRxiv version 2 · February 12, 2025