Foundation models to bridge the data scarcity and explainability gap in pancreatic cancer diagnosis.
Abstract
e16003 Background: Deep learning (DL) for medical imaging has traditionally emphasized standard classification metrics, such as accuracy, sensitivity, and specificity. However, performance gains alone are not sufficient for high variability scenarios such as pancreatic cancer diagnosis using endoscopy ultrasound (EUS). In these settings, clinical confidence requires decisions grounded in anatomically and clinically meaningful image regions. Furthermore, AI-ready EUS data is limited. To mitigate this, we compared conventional Convolutional Neural Networks (CNNs) against a vision transformer-based Ultrasound Foundation Model (USFM) with minimal fine-tuning on a small EUS dataset. Methods: We used an open EUS dataset of 50 subjects (18 pancreatic ductal adenocarcinoma cases and 32 healthy controls). USFM along with two CNN models, ResNet-50 (RN50) and EfficientNet-V2 (EN2), were fine-tuned for classification of cancer vs normal and tested with subject-wise 5-fold cross-validation (CV). Specifically, we fine-tuned the USFM by appending a two-layer classifier decoder to the publicly available encoder and utilized the last attention layer for generating attention maps. We also implemented a quantitative explainability framework based on the spatial overlap between the ground-truth lesion and AI model saliency maps (i.e., Attention Maps for USFM and Gradient-weighted Class Activation Maps for RN50 and EN2) using the Saliency-to-Segmentation Area under the Curve (AUC) metric. Results: USFM had the best performance and stability. It achieved the highest Accuracy (92%) and PPV (0.90), surpassing the best supervised baseline EN2 (Accuracy: 90%; PPV: 0.87). Furthermore, the alignment between model saliency and ground truth improved substantially. USFM's attention maps achieved a mean Saliency-to-Segmentation AUC of 0.85, outperforming the GradCAM results of EN2 (0.70) and RN50 (0.61). Qualitative and quantitative analyses showed that, unlike conventional DL models, transformer-based USFM can consistently concentrate on lesion regions. Conclusions: Results show that USFM, once finetuned with a small amount of downstream data, can improve predictive performance and interpretability. Validated through a 5-fold CV, the USFM not only achieved superior classification metrics with significantly fewer training epochs but also demonstrated substantially better localization of pathological regions in saliency maps. These findings suggest that FMs offer a more explainable method, effectively mitigating constraint of small data sizes while increasing reliability by ensuring decisions are grounded in anatomically and clinically relevant features. Performance comparison of EfficientNetV2, ResNet50, and USFM. Model Acc PPV NPV Sensitivity F1 Score Saliency to Segmentation AUC Score EfficientNetV2 90% 0.87 0.94 0.86 0.85 0.70 ResNet50 87% 0.86 0.89 0.77 0.80 0.61 USFM 92% 0.9 0.93 0.86 0.87 0.85
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (9)
Samin Yaser
University of Central Florida, Orlando, FL
Mahad Ali
University of Central Florida, Orlando, FL
M Iffat Hossain
University of Central Florida, Orlando, FL
Farhan Fuad Abir
University of Central Florida, Orlando, FL
Bryan Allinson
Vanquish Bio, Longwood, FL
Anthony Mango
Orlando Health, Orlando, FL
Juan Pablo Arnoletti
Orlando Health, Orlando, FL
Charles Robertson
Orlando Health, Orlando, FL
Laura Brattain
University of Central Florida, Orlando, FL