Concordance of response-based clinical trial and machine learning–generated real-world end points.
Abstract
e13625 Background: Real-world evidence (RWE) is increasingly used to complement clinical trial data in oncology, providing rapid insights to inform study design and drug development. Using deep learning, natural language processing (NLP)-based machine learning models, we developed a real-world response (rwR) approach.This study evaluates the concordance between clinical trial and real-world (rw) end points in patients with stage IV non–small cell lung cancer (NSCLC) treated with first-line platinum plus pemetrexed chemotherapy. Methods: This retrospective study compared response-based outcomes generated from patients included in the control arm of IMpower132 with a trial-aligned cohort of rw patients selected from the US-nationwide Flatiron Health electronic health record (EHR)-derived deidentified database. Rw patients were aligned to key trial inclusion/exclusion criteria and further adjusted using propensity score weighting on selected baseline characteristics including demographics (eg, age, race) and clinical factors (eg, Eastern Cooperative Oncology Group [ECOG] performance status, metastatic sites). rwR was generated using NLP-based machine learning models trained on expert human-abstracted data (training set N ~12 000 patients) to extract clinician-documented change in disease burden (ie, complete response, partial response, stable disease, progressive disease, unknown) at each imaging-based disease assessment timepoint. Trial response data were captured according to a RECIST-based trial protocol. End points included response rates (rwRR vs objective response rate [ORR]), duration of response (rwDOR vs DOR), and progression-free survival (rwPFS vs PFS). Concordance was evaluated using logistic regression for response rates and Cox regression for DOR and PFS. Results: The rw cohort (N = 494) was well aligned with the clinical trial cohort (N = 275) after weighting, with standardized mean differences below 0.1 across all selected baseline characteristics. The rwRR was 34%, compared with the 38% ORR observed in the trial cohort (OR, 0.83 [95% CI, 0.54-1.27]). The median rwDOR was 5.7 months (95% CI, 4.3-7.7), compared with 6.9 months (95% CI, 4.5-8.3) for DOR in the trial cohort (HR, 1.18 [95% CI, 0.86-1.62]). rwPFS was 5.5 months (95% CI, 4.5-7.1), closely aligned with the 5.4 months (95% CI, 4.3-5.7) observed in the trial cohort (HR, 0.96 [95% CI, 0.79-1.16]). Conclusions: NLP-based ML models enable scalable and reliable generation of response-based end points from EHRs. Concordance between trial and rw end points highlights the utility of ML-driven approaches to advance RWE in oncology, particularly when leveraging large-scale, clinically rich, and well-curated training datasets.
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (11)
Qianyi Zhang
State Key Laboratory of Metal Matrix Composites School of Materials Science and Engineering Shanghai Jiao Tong University Shanghai P. R. China
Konstantin Krismer
Flatiron Health, New York, NY
Yichen Lu
Qianyu Yuan
Flatiron Health, New York, NY
Aaron Dolor
Flatiron Health, New York, NY
Auriane Blarre
Flatiron Health, New York, NY
Aaron B. Cohen
Flatiron Health, New York, NY
Tori Williams
Flatiron Health, New York, NY
Sophia Maund
Genentech, Inc., South San Francisco, CA
Minu K Srivastava
Genentech, Inc., South San Francisco, CA
Kelly Magee
Flatiron Health, New York, NY