Beyond cellular locality: Vision transformer–based modeling for diagnosis and classification of blood cancers from peripheral blood smears.
Abstract
e18562 Background: Accurate morphologic assessment of peripheral blood smears remains central to hematologic malignancy diagnosis but is time-intensive and subject to interobserver variability, particularly in early or morphologically ambiguous disease. Convolutional neural networks (CNNs) have demonstrated utility in hematologic image analysis; however, their reliance on localized receptive fields may limit modeling of global cellular context and spatial relationships. Vision Transformers (ViTs) leverage self-attention mechanisms to enable long-range contextual reasoning across entire images. We evaluated the performance of a transformer-based architecture for automated classification of hematologic malignancies from peripheral blood smear images. Methods: This retrospective diagnostic modeling study utilized publicly available, de-identified peripheral blood smear image datasets derived from ALL-IDB–based repositories, including acute lymphoblastic leukemia, additional leukemic subtypes, and normal hematologic controls annotated by expert hematopathologists. Images were standardized, augmented, and divided into training and validation cohorts using stratified sampling. A pretrained Vision Transformer B/16 architecture was fine-tuned for multi-class classification. Input images (224×224 pixels) were partitioned into 16×16 patches and processed through 12 transformer encoder layers with multi-head self-attention. Model performance was assessed using accuracy, sensitivity, specificity, F1 score, and area under the receiver operating characteristic curve (AUROC). Performance was compared with published CNN benchmark results on similar datasets. Results: The Vision Transformer demonstrated high diagnostic performance across hematologic categories, achieving overall classification accuracy exceeding 99% with AUROC values greater than 0.99. Attention-based global modeling improved discrimination of leukemic blast populations with overlapping cytomorphologic features and reduced misclassification in visually heterogeneous samples. Performance remained stable across variable staining conditions and cell population distributions and was comparable to or exceeded reported CNN-based benchmarks on ALL-IDB–derived datasets. Conclusions: Transformer-based global attention modeling enables accurate and robust classification of hematologic malignancies by capturing spatial and population-level morphologic context beyond conventional localized feature extraction. These findings support the potential role of Vision Transformer architectures as AI-assisted diagnostic decision-support tools to augment hematopathology workflows. Prospective validation using real-world clinical datasets is warranted to assess generalizability and clinical integration.
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (10)
Nayanika Chowdary Tummala
NYMC at St. Mary’s General Hospital and Saint Clare’s Health, Denville, NJ
Elangovan Krishnan
AIM DOCTOR, Thiruvallur, India, India
Shankar Biswas
Jansi Rani Sethuraj
AIM DOCTOR, Thiruverkadu, India
Kavin Elangovan
AIM DOCTOR, Houston, Texas, United States
Ramya Elangovan
AIM DOCTOR, Houston, Texas, United States
Gowrishankar Palaniswamy
8Medical University of South Carolina, Lancaster, United States
Sophia Ahmed
Hammad Khan
Ayub Medical College, Abbottabad, Pakistan
Sravani Bhavanam
2Brookdale University Hospital and Medical center, Brooklyn, United States