Improving automated prostate cancer risk assessment: An ensemble LLM framework for clinical T-stage inference and NCCN risk stratification.
Abstract
e17005 Background: We previously developed a hybrid framework combining LLMs and rule-based algorithms (RBA) for automated risk stratification in localized prostate cancer (locPCa), achieving 89% accuracy. However, clinical T-stage (cTS) inference from narrative MRI reports remained the primary source of errors. Herein, we evaluated whether ensemble approaches using multiple state-of-the-art LLMs could improve cTS inference and overall NCCN risk stratification accuracy. Methods: This study included patients with locPCa presenting with at least one positive prostate biopsy and MRI report available. Building on our previously validated hybrid framework that extracts structured phenotypic variables (PSA, prostate volume, cTS, Gleason patterns, grade group, positive and examined cores) from unstructured reports using LLMs and applies NCCN-concordant RBA for risk stratification, we focused on improving cTS inference. Four HIPAA-compliant LLMs (GPT-4.1, GPT-5.1, Gemini-2.5-Pro, Gemini-2.5-Flash) were evaluated using structured zero-shot prompts iteratively refined through oncologist feedback. Majority voting was employed to generate consensus cTS predictions. Performance was evaluated using weighted accuracy. Results: A total of 358 patients were included. The median age at diagnosis was 65 years (IQR: 60-69); the majority were White (92%) and non-Hispanic (94%). cTS distribution was T2a (51%), T2c (20%), T3a (16%), T3b (7%), T2b (4%), and T4 (2%). The most prevalent risk category was unfavorable intermediate (38%) followed by favorable intermediate (20%), high risk (19%), very high risk (16%), low (5%), and very low (1%). Individual LLM accuracy for cTS inference ranged from 85-90%, with Gemini-2.5-Pro achieving highest accuracy (90%), followed by GPT-4.1 (88%), GPT-5.1 (87%), and Gemini-2.5-Flash (85%). The four-model ensemble reached consensus in 92% of cases, achieving 93% cTS inference accuracy. When integrated into NCCN risk stratification schema, single-model approaches achieved 89-92% accuracy, while ensemble consensus cases achieved 95% accuracy. Conclusions: The ensemble framework using multiple LLMs significantly improves cTS inference and achieves 95% NCCN risk stratification accuracy in consensus cases (92% of patients), representing a 6-percentage-point improvement over our previous single-model framework (89%). This strategy demonstrates that ensemble methods can overcome individual model limitations for complex clinical tasks and holds promise for reliable AI-assisted risk stratification.
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (8)
Umair Ayub
1Mayo Clinic, Division of Hematology/Oncology, Department of Internal Medicine, Phoenix, United States
Syed Arsalan Ahmed Naqvi
Mayo Clinic, Phoenix, AZ
Muhammad Umar Afzal
Mayo Clinic Arizona, Scottsdale, AZ
Salman Ayub Jajja
NYMC-LANDMARK MEDICAL CENTER, RI, Woonsocket, Rhode Island, United States
Ji-Eun Irene Yum
Mayo Clinic Alix School of Medicine, Phoenix, AZ
Yousef Zakharia
Division of Hematology and Medical Oncology, Department of Internal Medicine Mayo Clinic Phoenix Arizona USA
Parminder Singh
Department of Medicine, Mayo Clinic Alix School of Medicine, Phoenix, AZ
Irbaz Bin Riaz
Irbaz Bin Riaz, MD, PhD; R. Bryan Rumble, MSc; Thomas A. Hope, MD; Giuseppe Procopio, MD; and Neha Vapiwala, MD; Mayo Clinic, Phoenix, AZ; American Society of Clinical Oncology, Alexandria, VA; University of California, San Francisco, San Francisco, CA; Fondazione IRCCS Istituto Nazionale dei Tumori di Milano, Milan, Italy; and University of Pennsylvania Abramson Cancer Center, Philadelphia, PA