Capsule networks for cancer imaging applications.
Abstract
e13687 Background: Computer vision models are increasingly used in clinical oncology, with several FDA-approved algorithms now available for diagnostic imaging and pathologic analysis. Ensuring safe clinical deployment of these models requires evaluation not only of their predictive performance but also the stability of their predictions. While convolutional neural networks (CNNs) and vision transformers (ViTs) are the most widely used computer vision models, their behavior has been shown to be unstable, with small pixel-level changes leading to incorrect predictions. This instability raises concerns, as noise or imaging artifacts could render models unreliable. Capsule Networks (CapsNets) offer unique advantages compared to CNNs and ViTs including compact model size and lower training data requirements, resulting in increased model stability. We evaluated the utility of CapsNets for oncology imaging applications by studying both predictive performance and model stability. Methods: Three oncologic datasets were selected for analysis, focused on lung CT nodule classification (n = 1,633), breast ultrasound malignancy detection (n = 780), and hematopathology classification (n = 17,092). For each task, we trained two CNNs (ResNet-18 and ResNet-50), one ViT (MedViT), and two CapsNets (DR-CapsNet and BP-CapsNet). Predictive performance was evaluated using area under the receiver operating characteristic curve (AUC). To test model stability, AUC was measured on slightly perturbed images generated by projected gradient descent. Latent space embeddings were inspected to identify underlying reasons for model stability. Results: All models demonstrated comparable baseline performance across all datasets, with AUCs ranging from 0.860-0.996 for CNNs, 0.889-0.998 for the ViT, and 0.856-0.993 for the CapsNets. CapsNets, however, exhibited greater model stability than the CNNs and ViT. On slightly perturbed images (ε = 0.032), CapsNets achieved stable AUCs ranging from 0.856-0.987 (1.65% average decline) across all datasets. The CNNs and ViT had considerably worse AUCs ranging from 0.305-0.712 (37.5% average decline) and 0.410-0.678 (36.9% average decline), respectively. Latent space analysis showed greater feature stability for CapsNets. Notably, CapsNets also achieved these gains with substantially smaller model sizes, using over 35% fewer parameters than all other models. Conclusions: CapsNets have significant utility in oncologic image classification. Compared to CNNs and ViTs, CapsNets consistently exhibited similar performance and superior model stability with the added benefit of decreased model size. Model Model Size (x 10 6 parameters) Baseline AUC Range Perturbed AUC Range Average % AUC Decline ResNet-18 11.2 0.892-0.996 0.559-0.712 29.8 ResNet-50 23.5 0.860-0.995 0.305-0.652 45.1 MedViT 31.1 0.889-0.998 0.410-0.678 36.9 DR-CapsNet 7.0 0.880-0.993 0.857-0.898 3.2 BP-CapsNet 7.0 0.856-0.989 0.856-0.987 0.1
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (7)
Anand Srinivasan
Department of Applied Mathematics and Theoretical Physics, Centre for Mathematical Sciences
Durga Vahini Sritharan
Yale School of Medicine, New Haven, CT
Saahil Chadha
Yale School of Medicine, Department of Therapeutic Radiology, New Haven, CT
Daniel Fu
Yale School of Medicine, Department of Therapeutic Radiology, New Haven, CT
Gregory Aaron Breuer
Yale School of Medicine, Department of Internal Medicine (Medical Oncology), New Haven, CT
Jahid Omar Hossain
Department of Therapeutic Radiology, Yale School of Medicine, New Haven, CT
Sanjay Aneja
Department of Therapeutic Radiology, Yale School of Medicine, New Haven, CT