Digital twin models for counterfactual prediction in NSCLC: A validation study using clinical trial data.
Abstract
e20517 Background: Digital twin models (DTMs) developed on real world data can estimate a patient’s response to treatments they did not receive. These counterfactual (CF) predictions can help reduce a trial’s sample size without compromising its statistical power. Validation of DTM-based CF predictions remains limited. We evaluated CF outcome predictions in NSCLC by benchmarking DTM estimates against observed outcomes from IMpower131 and 132. 1 We then assessed potential trial sample size reductions enabled by DTMs. Methods: We trained machine learning (ML) models (penalized logistic regression [LR], XGBoost [XGB], MLP, GAT) to predict real-world overall survival (rwOS) for patients with stage IV NSCLC initiating 1L platinum chemotherapy (chemo) 2011-2016 from the Flatiron Health Research Database. Features were generated using structured and LLM-extracted clinical details (e.g. Charlson comorbidity index [CCI]; sites of metastases [SOM]). Models were internally validated on a held-out test set and externally validated (MAD 2 , rwOS and hazard ratio [HR] comparisons) using patient-level data from IMpower131 and 132. Predictions were made for stage IV patients in 1) the control arm to compare model predictions with observed trial outcomes (Table, Control Arm Validation) and 2) the experimental arm to simulate their outcomes as if they had instead received control-arm therapy (Table, CF Treatment Effect). We estimated trial sample size reduction based on the C-index from Cox Proportional Hazards models with model-based prognostic scores for covariate adjustment. Results: Key features across models included ECOG status, albumin, Brain+liver SOM, and CCI. LR and XGB performed best for IMpower131 and 132 respectively, with MAD 2 < 5% and predicted (pred) control arm rwOS similar to those observed (obs) in the trials 3 . CF predictions yielded HRs consistent with trial results (Table). Estimated sample size reduction was 9-15% and 15-21% for IMpower131 and 132 respectively. Conclusions: Validated DTMs demonstrate the feasibility of accurately predicting CF outcomes with advanced ML. This approach can enable efficient, well-powered trials with smaller sample sizes and supports a path toward broader validation and regulatory adoption. IMpower 131 IMpower 132 Control Arm Validation Pred vs Obs Median rwOS (months) 10 vs 12.6 3 15 vs. 13.1 3 MAD 2 3.9% 1.8% C-Index 0.61 0.67 HR (Pred:Obs; 1 is perfect prediction) 0.97 [0.86-1.00] 4 1.03 [1.00-1.06] 4 CF Treatment Effect Obs trial HR 0.86 (0.73-1.02) 3 0.82 (0.66-1.01) 3 DTM-derived HR 0.81 (0.79-0.84) 4 0.85 (0.83-0.86) 4 1 IMpower131 & IMpower132: atezolizumab plus chemo vs chemo alone in squamous and non-squamous NSCLC respectively. 2 Mean absolute difference between pred & obs OS curves. 3 Observed results differ slightly from the published trials as this analysis restricted to subset of stage IV patients. 4 95% CI based on 1000 bootstrapped samples of trial data.
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (10)
Aaron B. Cohen
Flatiron Health, New York, NY
Sandra Griffith
Flatiron Health, New York, NY
Melissa Estevez
Flatiron Health, New York, NY
Brandon Arnieri
Genentech, South San Francisco, CA
Matthew H. Secrest
Genentech, South San Francisco, CA
Joseph Manfredonia
Flatiron Health, New York, NY
Marcello Ricottone
Flatiron Health, New York, NY
Richard Knoche
Flatiron Health, New York, NY
Jacqueline Law
Flatiron Health, New York, NY
Melina Elpi Marmarelis
Penn Medicine Abramson Cancer Center, Philadelphia, PA