Augmenting clinician workflow to enable delivery of guideline concordant cancer care with an LLM-based system.
Abstract
e13702 Background: Applying Large Language Models (LLMs) to multifactorial data for patients with cancer may facilitate guideline-concordant care delivery and ease clinician burden. However, LLM reasoning capabilities for oncology concepts in Electronic Health Records (EHRs) and complex Standards of Practice (SOPs) is unclear. Methods: We developed an expert-guided, transparent LLM system to aid in assessment of EHRs of newly diagnosed patients and extract clinical data necessary for SOP-guided decision-making in a real world medical setting (OpenAI GPT-4o, accessed 10/13 - 12/13/2024). To solve for known challenges of LLM-based clinical tools (hallucinations, SOP alignment, introspection, and real-world data integration), we built a 2-step system that (1) structures EHRs by identifying and extracting SOP-relevant information and (2) evaluates completion of each SOP-defined diagnostic step. We tested performance and efficiency of the final model in a retrospective cohort of 100 de-identified patients with colon or breast cancer. The LLM evaluated unstructured data available in patients’ EHRs through treatment initiation to assess their pre-treatment diagnostic workup. Accuracy and categorization of errors as well as time required to conduct reviews was also recorded. Results: Unstructured EHRs for colon cancer patients averaged 7061 words (median [mdn] 3806, 90th percentile [p90] 21,029). The LLM system distilled these into 36 distinct decision factors required to generate diagnostic recommendations. For breast cancer patients, data averaged 1433 words (mdn 324, p90 4287) and was distilled into 89 decision factors. During the extraction phase for both cancer types, common sources of error were age and artifacts of the de-identification process (redaction of medical information). For colon cancer, a mdn time of 2.3 minutes (p90 6.5, avg 3.5 min) was required to approve extracted decision factors (step 1) and a mdn time of 2.1 min (p90 6.9, avg 3.2 min) to finalize recommendations (step 2). For breast cancer, the mdn time for step 1 was 5.2 min (p90 16.5, avg 22.5 min) and the mdn time for step 2 was 2.1 min (p90 7.9, avg 15.5). Conclusions: This disease, SOP-specific, state-of-the-art LLM-driven approach demonstrates high fidelity information distillation, all with a total mdn clinician time of 6.2 min across two cancer types. The longer time-to-finalized recommendations for breast cancer cases compared to colon (mdn 7.1 vs 5.3 min) can be attributed to the breast cancer SOP having a greater number of decision factors requiring review during data extraction (89 vs 36). It is promising that despite the number of decision factors for breast cancer being > 2x those for colon cancer, the mdn time-to-finalized recommendations was only 34% higher. This showcases a path to improved efficiency and accuracy of cancer care delivery through continued refinement of the model and workflow integration.
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (11)
Divneet Mandair
University of Colorado Hospital, Aurora, CO
Travis Zack
1University of California San Francisco, Hematology and Oncology, San Francisco, United States
Allison W. Kurian
Stanford Cancer Institute, Stanford University School of Medicine, Stanford, CA
Keegan Duchicela
Color Health, Burlingame, CA
Anjali Zimmer
Color Health, Burlingame, CA
Wendy McKennon
Color Health, Burlingame, CA
Kiril Kafadarov
Color Health, Burlingame, CA
Munir Al - Dajani
Color Health, Burlingame, CA
Amina Lazrak
Color Health, Burlingame, CA
Rebecca A. Miksad
Color Health, Burlingame, CA
Othman Laraki
Color Health, Burlingame, CA