Evaluating the efficiency and quality of clinical summaries produced by a large language model (LLM).
Abstract
e13601 Background: The critical task of curating data from various sources into an accurate and concise summary can be tedious, time-consuming, and error-prone, making it a challenging aspect of an oncologist's workflow when preparing for consultations and tumor boards, and when making patient referrals. navify Clinical Hub for Oncology (navify CH) aggregates oncology-specific information from multiple data sources and summarizes it using large language model (LLM) technology. This study evaluated the efficiency and quality of clinical summaries prepared by navify CH compared with SimEPR, a simulated EHR environment with design patterns and workflows common across EHRs. Methods: Board-certified oncologists (n = 26) were recruited from four countries: United States (n = 8), United Kingdom (n = 9), Spain (n = 5), and Singapore (n = 4). Each oncologist was instructed to review 10 synthetic breast cancer cases and compile a comprehensive clinical summary to present to a tumor board (five cases using nCH, five using SimEPR). Cases varied in complexity, reflecting a mix of presentations for 1st, 2nd and 3rd line treatment. To minimize learning and sequencing bias, a counterbalanced order of systems was used with cases paired based on the quantity of clinical information to be summarized. Primary outcome measures included the time taken to complete each summary and the quality of each summary. Quality was assessed by an independent oncologist across three domains - completeness, correctness, and conciseness (via a 5-point Likert scale). Secondary measures included the acceptability of navify CH via a survey. Results: A total of 99 paired summaries were included. Participants took significantly less time (112s) to produce a summary using navify CH compared to SimEPR (415s vs 527s, p < 0.001). Summaries written using navify CH had a significantly higher completeness score (3.93 vs 3.13, p < 0.001). These LLM-aided summaries were as concise and correct as human-generated ones, with slight, non-significant improvements in conciseness (+0.09) and correctness (+0.02). Users found navify CH highly acceptable, with 88% recommending it to colleagues. Most participants (92%) felt nCH would save time, with 54% estimating savings of 5-10 minutes per patient. LLM summarization was reported as the most valued feature, with 96% recommending it to colleagues and 85% reporting they would prioritize it for immediate integration into their workflow. The majority (92%) of participants reported that it would be useful for tumor board and consultation preparation. Conclusions: This study validates the utility of LLM technology in clinical summarization to gain a comprehensive understanding of the patient's history. The potential of LLM technology to save time must be balanced with accuracy and trust. These findings support the real-world utility of navify CH in the oncology workflow across multiple geographies.
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (7)
Jack Halligan
Prova Health, London, United Kingdom
Mehul Patel
UNC - CHAPEL HILL, Chapel Hill, North Carolina, United States
Gareth Obery
Prova Health, London, United Kingdom
Nesrine Lajmi
Roche, Santa Clara, CA
Archana Dorge
Roche, Santa Clara, CA
Ashish Sharma
Ernest N. Lo
Roche, Santa Clara, CA