Comparative Evaluation of Twenty Vision Transformer Variants for Six-Class Gastrointestinal Endoscopic Image Classification
PDF File

Keywords

Gastrointestinal endoscopy
Hyperkvasir
Vision transformer
Image classification
Comparative evaluation

How to Cite

Comparative Evaluation of Twenty Vision Transformer Variants for Six-Class Gastrointestinal Endoscopic Image Classification. (2026). Computers and Electronics in Medicine, 3(2), 162-173. https://doi.org/10.69882/adba.cem.2026077

Abstract

This study compares 20 vision transformer variants for six-class gastrointestinal endoscopic image classification using a labeled subset of HyperKvasir. The evaluated classes were cecum, polyps, pylorus, retroflex-rectum, retroflex-stomach, and z-line. Repeated row totals from six confusion matrices indicated an evaluation set of 768 images. The experiment was configured for five-fold cross-validation with a maximum of 50 epochs and a random seed of 42; however, the numerical fold-level file contained results only for Fold 1. Learning curves supplied qualitative information for selected runs in Folds 2–5 and complete graphical coverage for EViT-B0, EViT-B1, and EViT-B3. In the available numerical results, MaxViT-Tiny achieved the highest accuracy (99.09%), F1 score (99.10%), precision (99.12%), and recall (99.09%), whereas TinyViT-21M achieved the highest AUC (99.98%). MaxViT had the highest mean family accuracy, while TinyViT had the highest mean family AUC. Accuracy and F1 were strongly correlated, but the association between accuracy and AUC was weak. The learning curves showed rapid convergence with model-dependent validation fluctuations, and the confusion matrices located most errors in cecum–polyp and polyp–retroflex-rectum pairs. These findings support further evaluation of hybrid local–global transformer designs. Complete fold-level results, patient-level split information, calibration analysis, and external validation are required before drawing conclusions about cross-fold or clinical generalizability.




PDF File

References

Alaca, Y. and B. Emin, 2024 Performance evaluation of hybrid approaches combining deep learning models and machine learning methods for medical kidney image classification. 9th Azerbaijan Congress on Life, Engineering, Mathematical, and Applied Sciences p. 3007.

Azad, R., A. Kazerouni, M. Heidari, E. K. Aghdam, A. Molaei, et al., 2024 Advances in medical image analysis with vision transformers: A comprehensive review. Medical Image Analysis 91: 103000.

Ba¸saran, E. and Y. Çelik, 2024 Skin cancer diagnosis using cnn features with genetic algorithm and particle swarm optimization methods. Transactions of the Institute of Measurement and Control 46: 2733–2744.

Borgli, H., V. Thambawita, P. H. Smedsrud, S. Hicks, D. Jha, et al., 2020 Hyperkvasir, a comprehensive multi-class image and video dataset for gastrointestinal endoscopy. Scientific Data 7: 283.

Cai, H., C. Li, M. Hu, C. Gan, and S. Han, 2023 Efficientvit: Lightweight multi-scale attention for high-resolution dense prediction. Proceedings of the IEEE/CVF International Conference on Computer Vision pp. 17302–17313.

Chicco, D. and G. Jurman, 2020 The advantages of the matthews correlation coefficient over f1 score and accuracy in binary classification evaluation. BMC Genomics 21: 6.

Cubuk, E. D., B. Zoph, J. Shlens, and Q. V. Le, 2020 Randaugment: Practical automated data augmentation with a reduced search space. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops pp. 3008–3017.

Deng, J., W. Dong, R. Socher, L.-J. Li, K. Li, et al., 2009 Imagenet: A large-scale hierarchical image database. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition pp. 248–255.

Dosovitskiy, A., L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, et al., 2021 An image is worth 16x16 words: Transformers for image recognition at scale. International Conference on Learning Representations.

Esteva, A., B. Kuprel, R. A. Novoa, J. Ko, S. M. Swetter, et al., 2017 Dermatologist-level classification of skin cancer with deep neural networks. Nature 542: 115–118.

Fawcett, T., 2006 An introduction to roc analysis. Pattern Recognition Letters 27: 861–874.

Guo, H., S. A. Somayajula, R. Hosseini, et al., 2024 Improving image classification of gastrointestinal endoscopy using curriculum self-supervised learning. Scientific Reports 14: 3641.

Hatamizadeh, A., Y. Tang, V. Nath, D. Yang, A. Myronenko, et al., 2022 Unetr: Transformers for 3d medical image segmentation. IEEE/CVF Winter Conference on Applications of Computer Vision pp. 574–584.

He, H. and E. A. Garcia, 2009 Learning from imbalanced data. IEEE Transactions on Knowledge and Data Engineering 21: 1263–1284.

He, K., X. Zhang, S. Ren, and J. Sun, 2016 Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition pp. 770–778.

Huang, G., Z. Liu, L. van der Maaten, and K. Q. Weinberger, 2017 Densely connected convolutional networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition pp. 4700–4708.

Jha, D., P. H. Smedsrud, M. A. Riegler, D. Johansen, T. De Lange, et al., 2020 Kvasir-seg: A segmented polyp dataset. International Conference on Multimedia Modeling pp. 451–462.

Kelly, C. J., A. Karthikesalingam, M. Suleyman, G. Corrado, and D. King, 2019 Key challenges for delivering clinical impact with artificial intelligence. BMC Medicine 17: 195.

Kingma, D. P. and J. Ba, 2015 Adam: A method for stochastic optimization. International Conference on Learning Representations.

LeCun, Y., Y. Bengio, and G. Hinton, 2015 Deep learning. Nature 521: 436–444.

Lin, T.-Y., P. Goyal, R. Girshick, K. He, and P. Dollár, 2017 Focal loss for dense object detection. Proceedings of the IEEE International Conference on Computer Vision pp. 2980–2988.

Litjens, G., T. Kooi, B. E. Bejnordi, A. A. A. Setio, F. Ciompi, et al., 2017 A survey on deep learning in medical image analysis. Medical Image Analysis 42: 60–88.

Liu, X., S. C. Rivera, D. Moher, M. J. Calvert, and A. K. Denniston, 2020 Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: The CONSORT-AI extension. Nature Medicine 26: 1364–1374.

Liu, Z., Y. Lin, Y. Cao, H. Hu, Y. Wei, et al., 2021 Swin transformer: Hierarchical vision transformer using shifted windows. Proceedings of the IEEE/CVF International Conference on Computer Vision pp. 10012–10022.

Loshchilov, I. and F. Hutter, 2019 Decoupled weight decay regularization. International Conference on Learning Representations.

Lundberg, S. M., G. Erion, H. Chen, A. DeGrave, S. M. Pruthi, et al., 2020 From local explanations to global understanding with explainable AI for trees. Nature Machine Intelligence 2: 56–67.

Mehta, S. and M. Rastegari, 2022 Separable self-attention for mobile vision transformers. arXiv preprint arXiv:2206.02680.

Mongan, J., L. Moy, and J. Kahn, Charles E., 2020 Checklist for artificial intelligence in medical imaging (CLAIM): A guide for authors and reviewers. Radiology: Artificial Intelligence 2: e200029.

Mukhtorov, D., M. Rakhmonova, S. Muksimova, and Y.-I. Cho, 2023 Endoscopic image classification based on explainable deep learning. Sensors 23: 3176.

Pan, S. J. and Q. Yang, 2010 A survey on transfer learning. IEEE Transactions on Knowledge and Data Engineering 22: 1345–1359.

Pogorelov, K., K. R. Randel, C. Griwodz, S. L. Eskeland, T. de Lange, et al., 2017 Kvasir: A multi-class image dataset for computer aided gastrointestinal disease detection. Proceedings of the 8th ACM Multimedia Systems Conference pp. 164–169.

Ribeiro, M. T., S. Singh, and C. Guestrin, 2016 Why should I trust you?: Explaining the predictions of any classifier. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining pp. 1135–1144.

Roberts, M., D. Driggs, M. Thorpe, M. van der Schaar, C.-B. Schonlieb, et al., 2021 Common pitfalls and recommendations for using machine learning to detect and prognosticate for COVID-19 using chest radiographs and CT scans. Nature Machine Intelligence 3: 199–217.

Ronneberger, O., P. Fischer, and T. Brox, 2015 U-Net: Convolutional networks for biomedical image segmentation. Medical Image Computing and Computer-Assisted Intervention – MICCAI 2015 9351: 234–241.

Saito, T. and M. Rehmsmeier, 2015 The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PLOS ONE 10: e0118432.

Selvaraju, R. R., M. Cogswell, A. Das, R. Vedantam, D. Parikh, et al., 2017 Grad-CAM: Visual explanations from deep networks via gradient-based localization. Proceedings of the IEEE International Conference on Computer Vision pp. 618–626.

Shamshad, F., S. Khan, S. W. Zamir, M. H. Khan, M. Hayat, et al., 2023 Transformers in medical imaging: A survey. Medical Image Analysis 88: 102802.

Shen, D., G. Wu, and H.-I. Suk, 2017 Deep learning in medical image analysis. Annual Review of Biomedical Engineering 19: 221–248.

Shin, H.-C., H. R. Roth, M. Gao, L. Lu, Z. Xu, et al., 2016 Deep convolutional neural networks for computer-aided detection: CNN architectures, dataset characteristics and transfer learning. IEEE Transactions on Medical Imaging 35: 1285–1298.

Shorten, C. and T. M. Khoshgoftaar, 2019 A survey on image data augmentation for deep learning. Journal of Big Data 6: 60.

Tajbakhsh, N., J. Y. Shin, S. R. Gurudu, R. T. Hurst, C. B. Kendall, et al., 2016 Convolutional neural networks for medical image analysis: Full training or fine tuning? IEEE Transactions on Medical Imaging 35: 1299–1312.

Tan, M. and Q. V. Le, 2019 EfficientNet: Rethinking model scaling for convolutional neural networks. Proceedings of Machine Learning Research 97: 6105–6114.

Topol, E. J., 2019 High-performance medicine: The convergence of human and artificial intelligence. Nature Medicine 25: 44–56.

Touvron, H., M. Cord, M. Douze, F. Massa, A. Sablayrolles, et al., 2021 Training data-efficient image transformers and distillation through attention. Proceedings of Machine Learning Research 139: 10347–10357.

Touvron, H., M. Cord, and H. Jégou, 2022 DeiT III: Revenge of the ViT. European Conference on Computer Vision pp. 516–533.

Tu, Z., H. Talebi, H. Zhang, F. Yang, M. Peyfe, et al., 2022 MaxViT: Multi-axis vision transformer. European Conference on Computer Vision pp. 459–477.

Varma, S. and R. Simon, 2006 Bias in error estimation when using cross-validation for model selection. BMC Bioinformatics 7: 91.

Varoquaux, G., 2018 Cross-validation failure: Small sample sizes lead to large error bars. NeuroImage 180: 68–77.

Vaswani, A., N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, et al., 2017 Attention Is All You Need. Advances in Neural Information Processing Systems 30: 5998–6008.

Wiens, J., S. Saria, M. Sendak, M. Ghassemi, V. Liu, et al., 2019 Do no harm: A roadmap for responsible machine learning for health care. Nature Medicine 25: 1337–1340.

Wu, K., J. Zhang, H. Peng, M. Liu, B. Xiao, et al., 2022 TinyViT: Fast pretraining distillation for small vision transformers. European Conference on Computer Vision pp. 68–85.

Zech, J. R., M. A. Badgeley, M. Liu, A. B. Costa, J. J. Titano, et al., 2018 Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs. PLOS Medicine 15: e1002683.

Creative Commons License

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.