Accuracy of large language models for matching cancer patients to biomarker-driven clinical trials based on molecular profiles.
Abstract
e13627 Background: Clinical trials are an important treatment option for cancer patients, in particular if standard treatment modalities have been exhausted. With the availability of somatic genome profiling and targeted therapies an increasing number of biomarker-driven precision medicine trials are being conducted. It is a challenge even for experts to keep up with the number of relevant trials and to find all available options for a patient. In this study we explore whether large language models (LLM) could be used to augment the matching of trials to eligible patients based on their genomic profile and other information. Methods: We compiled a dataset of 678 full text descriptions of currently recruiting cancer clinical trials from clinicaltrials.gov and de-identified profiles from 100 patients with solid tumors. The profiles included basic demographics, a high-level diagnosis (e.g. lung adenocarcinoma) and somatic mutations from panel testing. We built an automated system to supply the trial and patient information to different LLMs, prompt the models to suggest suitable trials for each patient and automatically retrieve the results. We benchmarked the accuracy of the suggested trials with a manually curated ground truth dataset of 1107 positive patient-trial matches. Limitations of current LLMs are that they are not trained on real-time data (i.e. up-to-date clinical trial information) and have a limited context window size for providing trial data in the query. We therefore tried different strategies for condensing and pre-filtering the trial data: 1. Summarizing the full-text trial description into a single page using an LLM (chatGPT), 2. pre-filtering the trials by keyword search for the primary site of the patient (e.g. lung). Results: Of the LLMs we compared, Gemini had the highest accuracy across the 100 patient profiles (Sensitivity 45%, Specificity 22%). Without pre-filtering for primary site, sensitivity increased to 50% while specificity decreased to 16%, likely because keyword searching eliminates some valid trials with broad criteria (e.g. solid tumors). When providing trial data in full-text (not summarized), the accuracy decreased (Sn 37%, Sp 14% with pre-filtering, Sn 30%, Sp 5% without pre-filtering), suggesting that summarization is beneficial and too much information dilutes the matching. Common error modes included conflating lexicographically similar but clinically distinct entities, such as KRAS vs. NRAS and G12C vs. G12D, reflecting the probabilistic rather than exact matching behavior of LLMs. Conclusions: It can be assumed that patients will use available resources, including publicly available LLMs to search for treatment options. Our results show that LLM based approaches can find relevant potential trial options but are not comprehensive and also include a large number of false positives and therefore need to be interpreted with caution.
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (3)
Nourya A Cohen
Stanford University Department of Computer Science, Stanford, CA
Mohammad S. Esfahani
Henning Stehr