Predictive modeling by tumor type: Comparing performance and feature importance of machine learning–based predictive models in patients with hematologic and solid tumors.
Abstract
e13638 Background: Oncology practices increasingly rely on vast datasets from electronic health records and patient-reported outcomes (PROs). Prior work has shown that machine learning (ML)-based models can use these data to accurately predict adverse outcomes including acute care events (ACEs, i.e., emergency care and hospitalizations) and short-term mortality. However, it remains unknown whether these models perform differently among patients with hematologic versus solid tumors. Methods: We analyzed data from outpatient encounters at Massachusetts General Hospital Cancer Center from 9/2020–7/2022. Patients who had ACEs within 30 days were identified. A total of 193 variables were constructed based on vitals, labs, demographics, comorbidities, treatments, prior encounters and PROs. Patient encounters were subdivided into hematologic versus solid tumor cohorts. Univariate analyses were used to compare variables between encounters with and without ACEs. For the predictive analysis, each cohort was then divided into training (75%) and test (25%) sets. A variable screen filtered out the 133 most collinear variables, and 60 variables were ultimately used in each ML model. We tested multiple ML model architectures (xgboost, random forest, neural networks, logistic regression) and selected the best performing model for each cohort based on the area under the ROC curve (AUC) in the test sets. Variable importance scores were calculated for the best performing model in each cohort. Results: There were 13731 and 3681 encounters for patients with solid and hematologic tumors respectively, representing 3555 patients. Thirty-day ACEs occurred in 10.2% of encounters. In univariate analyses for solid patients, specific comorbidities (hypertension: 17.1% vs 10.0%, CAD 49.4% vs 36.9%, CKD 18.6% vs 10.9%; p < 0.001 for all) were most strongly associated with ACEs. Among hematologic patients, variables most strongly associated with ACEs were curative treatment intent therapy (25.8% vs 17.7 %, p < 0.001) and non-white race (12.5% vs 8.7%, p < 0.001). Random forest models performed best for solid and hematologic patients. Models had similar performance in solid ( AUC : 0.78; 95% CI 0.75-0.81) and hematologic patients ( AUC : 0.80; 95% CI 0.78-0.82). The most important variables for models with solid tumors were appetite, fatigue, and fever PROs, and for models with hematologic tumors were cardiac comorbidities, albumin, and hemoglobin. Conclusions: ML-based predictive models perform comparably well in predicting ACEs for patients with solid and hematologic malignancies. However, patients with solid and hematologic malignancies exhibit distinct clinical factors associated with ACEs. Predictive models for patients with cancer may perform best when models are developed on specific patient populations.
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (8)
Eashwar Somasundaram
Massachusetts General Hospital, Boston, MA
Patrick Connor Johnson
2Massachusetts General Hospital Cancer Center, Center for Lymphoma, Boston, United States
Areej El-Jawahri
1Cellular Immunotherapy Program, Massachusetts General Hospital Cancer Center, Harvard Medical School, Boston, MA
Andrea Pusic
Brigham and Women's Hospital, Boston, MA
Jason B. Liu
Brigham and Women's Hospital, Boston, MA
Therese Marie Mulvey
Massachusetts General Hospital Cancer Center, Boston, MA
Joseph Greer
1Mass General Brigham, Department of Psychiatry, Boston, United States
Thomas J Roberts
Massachusetts General Hospital, Harvard Medical School, Boston, MA