Tumor board–based multi-agent LLMs with guideline retrieval and consensus deliberation: Implications for hepatocellular carcinoma clinical reasoning.
Abstract
e16015 Background: Hepatocellular carcinoma (HCC), requires multidisciplinary decision-making across specialities. Multidisciplinary team (MDT) care has been associated with improved outcomes in HCC cohorts, yet major care-process gaps persist. Conventional single-call large language model (LLM) inference for clinical decision support may be constrained by limited guideline grounding. We evaluated whether an MDT-analogous multi-agent consensus architecture augmented with guideline retrieval improves LLM accuracy on validated HCC clinical reasoning tasks versus standard single-call inference. Methods: We developed a multi-agent system comprising three specialist agents (hepatology, oncology, radiology personas) and a supervisor agent. The system used retrieval-augmented generation (RAG) over a curated guideline corpus (AASLD, EASL, ESMO, BSG) implemented in a vector store, retrieving the top-5 passages per case. Performance was benchmarked against baseline single-call inference on a validated 88-item HCC Script Concordance Test (SCT) validated by 3 hepatologists with more than 10 years of experience. We evaluated a top large model (gpt-5.2) and a top small model (qwen3_vl_30b_thinking) selected from an initial benchmarking of 10 LLMs. Results: Baseline single-call accuracy ranged from 61.8% to 74.9% with no significant heterogeneity (Friedman p = 0.44). Large versus small model classes showed no significant difference (69.9% vs 67.4%; Wilcoxon p = 0.26). The multi-agent consensus architecture produced significant improvements for both model classes. For the large model (gpt-5.2), accuracy improved from 74.2% to 80.3% (+6.1 percentage points; 95% CI, 75.1-85.4%; p = 0.020). For the small model (qwen3-vl-30b), accuracy improved from 55.4% to 64.8% (+9.4; 95% CI, 57.0-72.5%; p = 0.029). Notably, the small model demonstrated the largest absolute gain; item-level analysis showed 17 previously incorrect items converted to correct versus only 3 errors introduced (McNemar p = 0.004). Run-to-run consistency was maintained or improved (gpt-5.2: 93%; qwen3-vl-30b: 55% to 73%). Conclusions: A tumor board-inspired multi-agent LLM architecture with guideline retrieval significantly improved HCC clinical reasoning accuracy for both large and small models, suggesting clinical relevance. The larger proportional benefit in smaller models raises the possibility of cost-effective deployment using open-weight architectures. Prospective validation should assess impact on tumor board efficiency, guideline adherence rates, and patient outcomes. Summary comparison table. Metric GPT-5.2 Solo GPT-5.2 Consensus Qwen Solo Qwen Consensus Accuracy 74.2% 80.3% 55.4% 64.8% 95% CI 68.1-79.8% 75.1-85.4% 48.8-62.0% 57.0-72.5% Wilcoxon p - 0.020 - 0.029 McNemar p - 0.19 - 0.004 Consistency 95.8% 93.0% 54.9% 73.2%
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (9)
Ernest Saenz
Division of Hepatology, Division of Clinical Bioinformatics, Department of Medicine II, Medical Faculty Mannheim, Heidelberg University, Mannheim, Germany
Santiago Rodriguez-Mora
Department of Internal Medicine, Division of Digestive Diseases, University of Cincinnati, Cincinnati, OH
Jimmy Daza
Division of Hepatology, Division of Clinical Bioinformatics, Department of Medicine II, Medical Faculty Mannheim, Heidelberg University, Mannheim, Germany
Santiago Arenas
Department of Internal Medicine, Division of Digestive Diseases, University of Cincinnati, Cincinnati, OH
Maria Fernanda Saavedra-Chacon
CES Clinic, Medellin, Colombia
Yeinis Paola Paola Espinoza-Herrera
Medical Faculty, Antioquia University, Medellin, Colombia
Juan Turnes
University Hospital Complex of Pontevedra, Galicia Sur Health Research Institute, Pontevedra, Spain
Andres Gomez-Aldana
College of Medicine, University of Cincinnati, Cincinnati, OH
Andreas Teufel
Division of Hepatology, Division of Clinical Bioinformatics, Department of Medicine II, Medical Faculty Mannheim, Heidelberg University, Mannheim, Germany