Literature Review | ProteinTalks: Enabling AI to Not Only “Recognize” Cells but Also “Predict” Drug Efficacy and Combination Therapies [Nature]

2026-09-24

IMG_256

 

The team led by Tiannan Guo at Westlake University, in collaboration with DP Technology, Peking University, and other institutions, has constructed the largest proteomics perturbation dataset to date and developed ProteinTalks, the first AI‑based foundation model grounded in protein network dynamics. This model demonstrates robust generalization across drug‑response prediction, combination‑therapy synergy screening, resistance‑mechanism elucidation, and clinical prognosis stratification, paving the way for a new paradigm in “virtual cell” modeling and precision drug discovery.

 

I. Research Background

Why build a “virtual cell”?

 

A virtual cell is a computational model of a real cell that can predict cellular life processes and dynamic behaviors on a computer. Constructing a reliable virtual cell requires comprehensive, time-resolved measurements of all intracellular components—while protein networks are central to understanding cellular biology and disease mechanisms.

 

However, compared with transcriptomic data, large-scale proteomic datasets—particularly perturbation data that incorporate dynamic temporal information—remain exceedingly scarce. In drug discovery, targeting a single protein without considering its contextual relationships within the protein network often yields limited efficacy or unexpected adverse effects. Consequently, systematically analyzing and modeling the dynamic changes of the proteome is essential for identifying novel drug targets, designing more precise therapeutic strategies, and ultimately constructing comprehensive virtual cell models.

 

In recent years, advances in data-independent acquisition mass spectrometry (DIA‑MS) have enabled high‑throughput proteomics, paving the way for leveraging large‑scale pretraining approaches to characterize protein dynamic landscapes. Against this technological backdrop, the present study has taken a pivotal step forward.

 

II. Data Scale

The largest proteomics dataset to date

The researchers used breast cancer cell lines as the experimental model and established a multidimensional perturbation proteomics experimental platform:

IMG_257

 

IMG_258

 

Quality control analyses revealed high consistency between biological and technical replicates (median Pearson correlation coefficient r = 0.95, median CV 0.15–0.16), confirming the high reliability of the data. This dataset, named the ProteinTalks Datasets (PTDS), is now publicly available at db.prottalks.com.

 

III. Model Architecture

ProteinTalks: The First Fundamental Model of Proteome Dynamics

Based on the aforementioned large-scale dataset, the research team developed a dynamics‑based foundational model called ProteinTalks. The model’s key innovation lies in integrating Neural Ordinary Differential Equations (Neural ODEs) with perturbation‑aware neural networks, thereby explicitly incorporating the temporal dimension into the architecture of the foundational model for the first time.

 

ProteinTalks Dual-Module Architecture

1) Kinetics Prediction Module (Module-1)

Input: undisturbed proteome + drug target information → Encoder → Neural ODE (using the RK4 solver) → Prediction of the perturbed proteome at 6 h / 24 h / 48 h → Computation of MSE loss (Loss₁) relative to experimental data. This module is responsible for learning the dynamic patterns of protein–protein network changes.

 

2) Drug Efficacy Prediction Module (Module-2)

Input: predicted perturbed proteome + drug molecular fingerprint (881-dimensional DMF) + drug physicochemical properties (55-dimensional DDP) + drug target → MLP → prediction of drug efficacy and combination synergy → BCE loss (Loss₂). This module is responsible for identifying core proteins associated with drug response.

 

3) Multi-task learning strategy

The weights of Loss₁ and Loss₂ (λ = 0.8) are dynamically adjusted using gradient cosine similarity, ensuring that the drug‑efficacy prediction task is prioritized when gradient directions conflict, thereby mitigating gradient competition between tasks.

IMG_259

 

Key differences from existing models: AlphaFold addresses the problem of protein structure prediction, while Geneformer and scGPT are based on single-cell transcriptomic data. ProteinTalks is the first foundational model to integrate large-scale proteomic perturbation data with temporal dynamics, filling a critical gap in predicting protein network dynamics.

 

IV. Prediction of Drug Efficacy

Outperforming the predictive capabilities of traditional machine learning models

ProteinTalks demonstrates significantly superior performance to conventional machine learning models on drug‑response prediction tasks: it outperforms transcriptomic baseline models across held-out scenarios, across cell lines, across drugs, in the untrained‑new‑drug PTNC setting, and across cancer types in the PTPC setting.

  • Same-distribution held-out : Accuracy 0.86, AUROC 0.92, AUPRC 0.84.
  • A cell line that is retained : On 17 breast cancer cell lines, the AUPRC was 0.89 and the AUROC was 0.95, with the transcriptomic model clearly outperforming others.
  • Leave one medication : There is still positive transfer to MOAs that were not encountered during training.
  • Completely untrained new drug PTNC : 81 new anticancer drugs achieved an AUPRC of 0.94, an AUROC of 0.89, and an accuracy of 0.85, whereas Geneformer’s performance dropped almost to the level of random guessing.
  • PTPC across cancer types : Melanoma, colorectal cancer, lung cancer, and pancreatic cancer all outperform all transcriptomic baselines even after few-shot learning.

IMG_260

V. Prediction of Drug Interactions

In drug‑combination prediction, ProteinTalks integrates the molecular fingerprints of two drugs with protein‑dynamic contextual information, successfully distinguishing synergistic from non‑synergistic pairs. It also guided wet‑lab validation, confirming strong synergistic effects—such as those of bosutinib plus tucatinib and bosutinib plus abemaciclib—in TNBC cell lines, thereby transforming a “virtual cell model” into an operational tool for drug screening, hypothesis generation, and precision‑medicine recommendations.

IMG_261

 

VI. Explainability

SHAP values reveal key proteins associated with drug resistance.

The interpretability of ProteinTalks is one of its key strengths. By computing SHAP (SHapley Additive exPlanations) values, researchers can quantify the contribution of each protein to drug‑sensitivity predictions, thereby identifying core proteins associated with drug resistance.

 

Interpretation at the channel level

Drugs with similar mechanisms of action (MOA) exhibit comparable patterns on SHAP pathway heatmaps. For example, the DNA repair pathway is critical for predicting sensitivity to alkylating agents; the PI3K–AKT–mTOR pathway lies at the heart of responses to PI3K/AKT inhibitors; and the estrogen‑response pathway predominates in the context of aromatase inhibitors.

IMG_262

 

Experimental validation of key drug-resistant proteins

Key proteins of AKR1C3 hormone‑related drugs and kinase inhibitors

One of the proteins with the highest SHAP values. The researchers used small interfering RNA (siRNA) to knock down AKR1C3 in BT20 cells and assessed its impact on docetaxel sensitivity using cytotoxicity assays. Knockdown efficiency was confirmed by mass spectrometry. Cytotoxicity assays demonstrated that AKR1C3 depletion enhanced the sensitivity of BT20 cells to docetaxel (Δln[IC50] = 4.24), supporting AKR1C3 as a protein prioritized by ProteinTalks and associated with docetaxel response.

 

Meaning ProteinTalks not only predicts “whether a drug is effective,” but also elucidates “why it is or is not effective”—by leveraging SHAP analysis to pinpoint specific proteins and pathways, thereby providing precise targets for subsequent experimental validation and mechanistic studies.

 

VII. Clinical Translation

From Cell Lines to Patients: Toward Precision Medicine

Transfer Learning: Cross-Modal Enhancement from Proteome to Transcriptome

 

The research team further explored the application of ProteinTalks in patient-derived xenograft (PDX) models. By pre-training on proteomic data and then fine-tuning on transcriptomic data using transfer learning, the model’s performance in predicting drug efficacy on PDX datasets was significantly improved:

IMG_263

VIII. Key Takeaways

What does this study bring?

1) Data scale breakthrough : The 38 million protein perturbation measurements constitute the largest proteomics perturbation dataset to date, laying a data foundation for AI-driven proteomics research.

2) The first dynamics-based model ProteinTalks is the first to explicitly incorporate temporal dynamics into a proteome‑based model, overcoming the limitation of previous models that relied solely on steady‑state data.

3) A New Paradigm in Drug Discovery The model demonstrates practical utility in predicting the efficacy of monotherapy (AUROC = 0.96), screening for synergistic drug combinations (with 4 out of 4 experiments successfully validated), and identifying drug‑resistance proteins.

4) Explainability Advantage SHAP analysis transforms models from “black boxes” into “transparent” ones, enabling precise identification of the key proteins and pathways that drive drug response and resistance.

5) Clinical Translation Pathway Through transfer learning, the proteome pre-training model can effectively enhance the predictive performance of transcriptomic data and has been validated to have prognostic stratification value in real-world patient cohorts.

6) An important step for the “virtual cell” This study demonstrates that a combination of limited large-scale perturbation measurements and dynamical AI models can be used to construct a virtual cellular modeling framework that spans both time and space.

 

IX. Future Prospects

As proteomic depth and throughput continue to improve, and PTM data are integrated, the ProteinTalks framework is poised to evolve into a more comprehensive “virtual cell” model, encompassing a broader range of disease types and biological processes.

 

Reference link:

https://www.nature.com/articles/s41586-026-11001-9