A large-scale, multi-target deep learning model for virtual genomic and molecular risk profiling in colorectal cancer.
Abstract
3526 Background: Molecular profiling of tumor biopsies is central to precision oncology, informing treatment selection, prognostication, and disease monitoring. Computational analysis of routine H&E slides offers a complementary, scalable approach that can accelerate biological insights and optimize downstream molecular testing. Recent advances in computational pathology have demonstrated that deep learning models can infer molecular features directly from digital H&E images, referred to as virtual genomics. However, most studies in colorectal cancer (CRC) have been constrained by limited cohort sizes, single-target predictions, and a lack of integration with longitudinal clinical outcomes. Here, we present a transformer-based deep learning framework designed for large-scale virtual genomic and molecular recurrence risk profiling in CRC. Methods: The model was trained on 45,155 patients with matched digital H&E images, whole-exome sequencing (WES), and longitudinal circulating tumor DNA (ctDNA) for molecular recurrence monitoring. We embedded all images using H-optimus-0 and trained a transformer-based multiple instance learning aggregation head for downstream virtual genomic predictions. We trained individual models to predict genes found in the MSK-IMPACT505 cancer gene panel and a single model that predicts all genes simultaneously. Lastly, we trained an independent model to predict the risk of molecular recurrence. All results were validated on an external cohort from The Cancer Genome Atlas (TCGA; n = 422 CRC patients). Results: Individual models predicted 379 mutated genes with an internal area under the receiver operating curve (AUROC) > 0.7. We validated 254 of these genes within TCGA with an AUROC > 0.7. Training a multi-task learning model to predict all genes simultaneously improved the AUROC for 85% of the genes. Based on NCCN guidelines in CRC, we further trained individual models to predict MSI status, BRAF V600E, KRAS G12D/V/G13D, and POLE/POLD1 exonuclease mutations with AUROCs of 0.96, 0.93, 0.84, and 0.86, respectively. Lastly, we developed an image-based molecular recurrence risk model trained directly on longitudinal ctDNA outcomes. The model stratified patients into low, medium, and high risk groups with a concordance index of 0.67. Compared to low-risk patients, the high-risk group exhibited a hazard ratio (HR) of 4.67, while the medium-risk group showed an HR of 1.98 (p-values < < 1E-5). Conclusions: This unified framework demonstrates that virtual genomics from routine H&E histopathology can enable scalable, cost-efficient inference of hundreds of clinically relevant genomic alterations and molecular recurrence risk in CRC. By leveraging universally available diagnostic slides, virtual genomics has the potential to expand access to precision oncology, optimize molecular testing strategies, and support earlier risk stratification.
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (10)
Erik N. Bergstrom
Tinghui Wu
Natera, Inc., Austin, TX
Michail Chatzianastasis
Natera, Inc., Austin, TX
Thinh Tran
Aaron M. Rosenfeld
Robert Burns
Natera, Inc., Austin, TX
Adham A. Jurdi
Natera, Inc., Austin, TX
Matthew Rabinowitz
MyOme, Inc, Menlo Park, California, United States
Helio Costa
Natera, Inc., Austin, TX
Frank Zhang
Natera, Inc., Austin, TX