A natural language processing algorithm to extract MRI and biopsy-related phenotypes for prostate cancer patients.
Abstract
356 Background: Extraction of data from unstructured medical records is a powerful but challenging tool for deriving novel information to drive research, improve clinical care and better inform guidelines. Here, we describe the creation and validation of a novel natural language processing (NLP) algorithm for extracting components of interest from biopsy and MRI reports for prostate cancer patients. Our algorithm specifically extracts the Gleason score from biopsy reports and maximum Prostate Imaging-Reporting and Data System (PI-RADS) score, Prostate-specific androgen (PSA) density, prostate volume, and prostate dimensions from MRI reports. Methods: MRI and biopsy pathology reports were extracted for a cohort of 155,570 patients diagnosed with prostate cancer between 1999 and 2024 either in the VA Cancer Registry System (VACRS) with prostate as their primary site of tumor or in the VA Corporate Data Warehouse with a relevant procedure or diagnosis code. These were annotated by a physician and a trained research scientist to generate data for the development and validation of our algorithm. Disagreements between annotators were adjudicated by a urologist. Our rule-based NLP algorithm was iteratively developed on 600 annotated and unannotated notes. The algorithm was validated on a manually annotated set of 250 MRI reports and 250 biopsy pathology notes from 378 patients at 78 VA centers for procedures between 2004 and 2024. Precision (true positives / (true positives + false positives)), recall (true positives / (true positives + false negatives)), and F1 score (2 * (precision * recall) / (precision + recall)) were computed to evaluate algorithm performance, with higher scores indicating better performance. Results: Our algorithm successfully extracted all five components from text reports with high precision and sensitivity. Item-level performance metrics for our algorithm on unseen data are reported in the table below. Post-hoc spot checking of algorithm performance on an additional sample of notes selected by procedure year revealed consistency of results across time for all components except prostate dimensions, for which a greater proportion of recent negative notes were false negatives. Further error analysis showed that many Gleason scores that our algorithm did not extract were repetitions of scores successfully extracted elsewhere in the note. Conclusions: Our NLP algorithm is able to derive structured data from medical reports with highly diverse formats with excellent accuracy and reliability. This approach has great potential to facilitate data extraction in a range of settings and to drive future research and clinical questions. Component Precision % Recall % F1 score % Gleason grade group 99.1 94.7 96.9 PI-RADS 95.5 92.7 94.1 PSA density 100.0 99.0 99.5 Prostate volume 96.9 94.5 95.7 Prostate dimensions 98.4 90.2 94.1
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (20)
John Culnan
2VA Boston Healthcare System, Boston, United States
Sergey Goryachev
Massachusetts Veterans Epidemiology Research and Information Center, Department of Veterans Affairs Healthcare System, Boston, MA
Daniel Chen
Oleg Soloviev
VA Boston Healthcare System, Boston, MA
Grace Lee
John Bihn
VA Boston Healthcare System, Boston, MA
June Corrigan
2VA Boston Healthcare System, Boston, United States
Karlynn N Dulberger
VA Boston Healthcare System, Boston, MA
Jennifer La
BOSTON UNIVERSITY SCHOOL MEDICINE, Boston, Massachusetts, United States
Kaitlin Swinnerton
Massachusetts Veterans Epidemiology Research and Information Center, Boston, MA
Tanya B. Dorff
Department of Medical Oncology and Therapeutics, City of Hope Comprehensive Cancer Center
Isla Garraway
UCLA David Geffen School of Medicine, Los Angeles, CA
Susan Halabi
From the Center for Cancer Research, National Cancer Institute, National Institutes of Health, Bethesda (A.B.A., N.S., S.N., L.L., L.C.), the Sidney Kimmel Comprehensive Cancer Center at Johns Hopkins, Baltimore (J.H.-C.), and the Investigational Drug Branch, Cancer Therapy Evaluation Program, National Cancer Institute, National Institutes of Health, Rockville (H.S., E.S.) — all in Maryland; the Alliance Statistics and Data Management Center, Mayo Clinic, Rochester, MN (K.V.B., M.O., C.M., G.P.B.); AdventHealth Cancer Institute and the University of Central Florida, Orlando (G.S.); Dana–Farber/Harvard Cancer Center, Boston (S.B., B.M.); UNC Lineberger Comprehensive Cancer Center, Chapel Hill (W.Y.K.), and Duke University Medical Center and Duke Cancer Institute, Durham (J.H., S.H.) — both in North Carolina; the University of Kansas Cancer Center, Westwood (R.P.); Memorial Sloan Kettering Cancer Center, New York (M.Y.T., M.J.M., J.E.R.), and Roswell Park Comprehensive Cancer Center, Buffalo (G.C.) — both in...
Stacy Loeb
NYU Langone Health, New York, NY
Nicholas George Nickols
Greater Los Angeles Department of Veterans Affairs Healthcare System, Los Angeles, CA
Matthew Rettig
Department of Medical Oncology, University of Southern California, Los Angeles, Los Angeles, CA
Martin W. Schoen
Division of Hematology and Medical Oncology, Department of Internal Medicine, Saint Louis University School of Medicine, St. Louis, MO
Channing Judith Paller
Sidney Kimmel Comprehensive Cancer Center at Johns Hopkins University School of Medicine, Baltimore, MD
Nathanael Fillmore
Massachusetts Veterans Epidemiology Research and Information Center, VA Boston Healthcare System, Boston, Massachusetts, United States
Matthew R. Cooperberg
University of California, San Francisco, San Francisco, CA