Tumor board–based multi-agent LLMs with guideline retrieval and consensus deliberation: Implications for hepatocellular carcinoma clinical reasoning.

E Ernest Saenz (Division of Hepatology, Division of Clinical Bioinformatics, Department of Medicine II, Medical Faculty Mannheim, Heidelberg University, Mannheim, Germany) S Santiago Rodriguez-Mora (Department of Internal Medicine, Division of Digestive Diseases, University of Cincinnati, Cincinnati, OH) J Jimmy Daza (Division of Hepatology, Division of Clinical Bioinformatics, Department of Medicine II, Medical Faculty Mannheim, Heidelberg University, Mannheim, Germany) S Santiago Arenas (Department of Internal Medicine, Division of Digestive Diseases, University of Cincinnati, Cincinnati, OH) M Maria Fernanda Saavedra-Chacon (CES Clinic, Medellin, Colombia) Y Yeinis Paola Paola Espinoza-Herrera (Medical Faculty, Antioquia University, Medellin, Colombia) J Juan Turnes (University Hospital Complex of Pontevedra, Galicia Sur Health Research Institute, Pontevedra, Spain) A Andres Gomez-Aldana (College of Medicine, University of Cincinnati, Cincinnati, OH) A Andreas Teufel (Division of Hepatology, Division of Clinical Bioinformatics, Department of Medicine II, Medical Faculty Mannheim, Heidelberg University, Mannheim, Germany)

Abstract

e16015 Background: Hepatocellular carcinoma (HCC), requires multidisciplinary decision-making across specialities. Multidisciplinary team (MDT) care has been associated with improved outcomes in HCC cohorts, yet major care-process gaps persist. Conventional single-call large language model (LLM) inference for clinical decision support may be constrained by limited guideline grounding. We evaluated whether an MDT-analogous multi-agent consensus architecture augmented with guideline retrieval improves LLM accuracy on validated HCC clinical reasoning tasks versus standard single-call inference. Methods: We developed a multi-agent system comprising three specialist agents (hepatology, oncology, radiology personas) and a supervisor agent. The system used retrieval-augmented generation (RAG) over a curated guideline corpus (AASLD, EASL, ESMO, BSG) implemented in a vector store, retrieving the top-5 passages per case. Performance was benchmarked against baseline single-call inference on a validated 88-item HCC Script Concordance Test (SCT) validated by 3 hepatologists with more than 10 years of experience. We evaluated a top large model (gpt-5.2) and a top small model (qwen3_vl_30b_thinking) selected from an initial benchmarking of 10 LLMs. Results: Baseline single-call accuracy ranged from 61.8% to 74.9% with no significant heterogeneity (Friedman p = 0.44). Large versus small model classes showed no significant difference (69.9% vs 67.4%; Wilcoxon p = 0.26). The multi-agent consensus architecture produced significant improvements for both model classes. For the large model (gpt-5.2), accuracy improved from 74.2% to 80.3% (+6.1 percentage points; 95% CI, 75.1-85.4%; p = 0.020). For the small model (qwen3-vl-30b), accuracy improved from 55.4% to 64.8% (+9.4; 95% CI, 57.0-72.5%; p = 0.029). Notably, the small model demonstrated the largest absolute gain; item-level analysis showed 17 previously incorrect items converted to correct versus only 3 errors introduced (McNemar p = 0.004). Run-to-run consistency was maintained or improved (gpt-5.2: 93%; qwen3-vl-30b: 55% to 73%). Conclusions: A tumor board-inspired multi-agent LLM architecture with guideline retrieval significantly improved HCC clinical reasoning accuracy for both large and small models, suggesting clinical relevance. The larger proportional benefit in smaller models raises the possibility of cost-effective deployment using open-weight architectures. Prospective validation should assess impact on tumor board efficiency, guideline adherence rates, and patient outcomes. Summary comparison table. Metric GPT-5.2 Solo GPT-5.2 Consensus Qwen Solo Qwen Consensus Accuracy 74.2% 80.3% 55.4% 64.8% 95% CI 68.1-79.8% 75.1-85.4% 48.8-62.0% 57.0-72.5% Wilcoxon p - 0.020 - 0.029 McNemar p - 0.19 - 0.004 Consistency 95.8% 93.0% 54.9% 73.2%

Article Details

Volume / Issue Vol. 44, Issue 16_suppl
Published June 01, 2026
ISSN 0732-183X
Publisher Lippincott Williams & Wilkins

Journal Info

Journal of Clinical Oncology

Lippincott Williams & Wilkins

ISSN: 0732-183X Health Sciences

Authors (9)

E

Ernest Saenz

Division of Hepatology, Division of Clinical Bioinformatics, Department of Medicine II, Medical Faculty Mannheim, Heidelberg University, Mannheim, Germany

S

Santiago Rodriguez-Mora

Department of Internal Medicine, Division of Digestive Diseases, University of Cincinnati, Cincinnati, OH

J

Jimmy Daza

Division of Hepatology, Division of Clinical Bioinformatics, Department of Medicine II, Medical Faculty Mannheim, Heidelberg University, Mannheim, Germany

S

Santiago Arenas

Department of Internal Medicine, Division of Digestive Diseases, University of Cincinnati, Cincinnati, OH

M

Maria Fernanda Saavedra-Chacon

CES Clinic, Medellin, Colombia

Y

Yeinis Paola Paola Espinoza-Herrera

Medical Faculty, Antioquia University, Medellin, Colombia

J

Juan Turnes

University Hospital Complex of Pontevedra, Galicia Sur Health Research Institute, Pontevedra, Spain

A

Andres Gomez-Aldana

College of Medicine, University of Cincinnati, Cincinnati, OH

A

Andreas Teufel

Division of Hepatology, Division of Clinical Bioinformatics, Department of Medicine II, Medical Faculty Mannheim, Heidelberg University, Mannheim, Germany