Assessment of pathologic responses from unstructured pathology reports using large language models.
Abstract
866 Background: Emerging data suggests that Language Models (LLMs) provide an unprecedented opportunity to extract critical data from the unstructured pathology reports. However, the performance of state of art (SOTA) models on clinically meaningful tasks such as assessment of pathologic responses in patients receiving neoadjuvant chemotherapy (NAC) in muscle invasive bladder cancer(MIBC) is not well known. Methods: This retrospective cohort included patients with pathologically confirmed Muscle-Invasive Bladder Cancer from the year 2018 to 2024. Selected patients included those who underwent Neoadjuvant chemotherapy followed by definitive cystectomy at Mayo Clinic. Gold standard labels were manually curated by trained clinicians. The state of the art (SOTA) LLM – GPT4 – was utilized using a structured zero-shot prompt to extract relevant phenotypic variables such as tumor site, tumor histology and additional histological variants if any, presence of in-situ / invasive carcinoma, presence of lamina propria / muscularis propria / lympho-vascular / perineural invasion, and T-stage from unstructured pathology reports. Structured prompts were iteratively developed for each data variable and validated using ~5% of the total dataset. The extracted variables were manually compared and labelled as True positive / True negative / False positive / False negative. Final performance was assessed against a held-out expert annotated test dataset using evaluation metric (accuracy). Results: A total of 200 reports from 99 patients were extracted by the LLM. Of the 200 reports, 41 duplicate / other pathology reports (e.g. – autopsy, cholecystectomy, etc.) were removed during final assessment. Significant characteristics such as tumor site, histology, and additional histology variant had an accuracy of 89%, 87%, and 91% respectively. Notably, accuracy of detection of in-situ versus invasive carcinoma was 70% versus 91%. Specific variables such as lamina propria / muscularis propria invasion / T-stage / lympho-vascular invasion / perineural invasion were answered with the accuracy of 75%, 75%, 67%, 90%, and 94% respectively. Conclusions: LLMs have the potential to extract critical data from unstructured pathology reports and automate assessment of pathologic responses in MIBC patients receiving NAC. Iterative improvements in performance can be achieved by improving prompt design using expert clinical input. This process can be potentially scaled to automate cancer registries.
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (7)
Swati Popli
Mayo Clinic in Arizona, Scottsdale, AZ
Muhammad Umair Anjum
The Wright Center for GME, Scranton, Pennsylvania, United States
Syed Arsalan Ahmed Naqvi
Mayo Clinic, Phoenix, AZ
Prateek Jain
Sinai Hospital Baltimore, Baltimore, MD
Parminder Singh
Department of Medicine, Mayo Clinic Alix School of Medicine, Phoenix, AZ
Yousef Zakharia
Division of Hematology and Medical Oncology, Department of Internal Medicine Mayo Clinic Phoenix Arizona USA
Irbaz Bin Riaz
Irbaz Bin Riaz, MD, PhD; R. Bryan Rumble, MSc; Thomas A. Hope, MD; Giuseppe Procopio, MD; and Neha Vapiwala, MD; Mayo Clinic, Phoenix, AZ; American Society of Clinical Oncology, Alexandria, VA; University of California, San Francisco, San Francisco, CA; Fondazione IRCCS Istituto Nazionale dei Tumori di Milano, Milan, Italy; and University of Pennsylvania Abramson Cancer Center, Philadelphia, PA