SELF-SUPERVISED VISION TRANSFORMERS FOR EARLY DISEASE DETECTION IN LOW-RESOURCE MEDICAL IMAGING ENVIRONMENTS FRAMEWORK
Main Article Content
Abstract
This study proposes a Self-Supervised Vision Transformer (ViT) framework for early disease detection in low-resource medical imaging environments. The research addresses major challenges in healthcare systems, including limited labeled medical imaging datasets, inadequate computational resources, and a shortage of medical experts. Traditional Convolutional Neural Network (CNN)-based models generally require large annotated datasets and extensive computational power, limiting their applicability in resource-constrained healthcare settings. To overcome these limitations, the proposed framework integrates self-supervised learning techniques with a Vision Transformer architecture to improve medical image classification performance using limited labeled data. The framework uses chest X-ray images for pneumonia detection and employs preprocessing methods, including normalization, augmentation, resizing, and denoising, to improve image quality and consistency. Self-supervised learning techniques, including contrastive learning and masked image modeling, are used to learn meaningful image representations from unlabeled data. The Vision Transformer architecture captures both local and global contextual information through self-attention mechanisms, enhancing feature extraction and disease representation. Experimental evaluation is conducted using metrics such as accuracy, precision, recall, F1-score, confusion matrix, and ROC-AUC, with a comparative analysis against the ResNet50 CNN model. The results demonstrate that the proposed Vision Transformer framework achieves strong disease-sensitive classification performance and excellent recall for pneumonia detection, while reducing dependency on large annotated datasets. Although the CNN model achieved higher overall classification accuracy, the Vision Transformer demonstrated superior contextual learning and disease-discrimination capabilities. The study highlights the potential of self-supervised Vision Transformers for scalable, efficient, and cost-effective AI-assisted medical image analysis in low-resource healthcare environments.
Downloads
Article Details
Issue
Section

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
How to Cite
References
1. Abdulrazzaq, M. M., Ramaha, N. T., Hameed, A. A., Salman, M., Yon, D. K., Fitriyani, N. L., ... & Lee, S. W. (2024). Consequential advancements of self-supervised learning (SSL) in deep learning contexts. Mathematics, 12(5), 758. https://doi.org/10.3390/math12050758 DOI: https://doi.org/10.3390/math12050758
2. Aburass, S., Dorgham, O., Al Shaqsi, J., Abu Rumman, M., & Al-Kadi, O. (2025). Vision Transformers in Medical Imaging: A Comprehensive Review of Advancements and Applications Across Multiple Diseases. Journal of Imaging Informatics in Medicine, 38(6), 3928–3971. https://doi.org/10.1007/s10278-025-01481-y DOI: https://doi.org/10.1007/s10278-025-01481-y
3. Chen, X., Xie, S., & He, K. (2021). An Empirical Study of Training Self-Supervised Vision Transformers (Version 4). arXiv. https://doi.org/10.48550/ARXIV.2104.02057 DOI: https://doi.org/10.1109/ICCV48922.2021.00950
4. Cleveland Clinic (2022). X-Ray: What It Is, Types, Preparation and Risks. [online] Cleveland Clinic. Available at: https://my.clevelandclinic.org/health/diagnostics/21818-x-ray.
5. Egbuna, I., Okei, N. C., Ajoku, E., Edobor, O., Issah, P., Olatokun, T., & Akinode, A. O. (2025). Advancing early disease detection with AI: innovations in medical imaging, EHR analytics, and wearable technologies. International Journal of Life Science Research Archive, 9, 120-140. https://doi.org/10.53771/ijlsra.2025.9.1.0049 DOI: https://doi.org/10.53771/ijlsra.2025.9.1.0049
6. Elyan, E., Vuttipittayamongkol, P., Johnston, P., Martin, K., McPherson, K., Moreno-García, C. F., ... & Sarker, M. M. K. (2022). Computer vision and machine learning for medical image analysis: recent advances, challenges, and way forward. Artificial Intelligence Surgery, 2(1), 24-45. http://dx.doi.org/10.20517/ais.2021.15 DOI: https://doi.org/10.20517/ais.2021.15
7. Huang, S.-C., Pareek, A., Jensen, M., Lungren, M. P., Yeung, S., & Chaudhari, A. S. (2023). Self-supervised learning for medical image classification: a systematic review and implementation guidelines. Npj Digital Medicine, 6(1). https://doi.org/10.1038/s41746-023-00811-0 DOI: https://doi.org/10.1038/s41746-023-00811-0
8. Jiang, J., Tyagi, N., Tringale, K., Crane, C., & Veeraraghavan, H. (2022). Self-supervised 3D Anatomy Segmentation Using Self-distilled Masked Image Transformer (SMIT). In Lecture Notes in Computer Science (pp. 556–566). Springer Nature Switzerland. https://doi.org/10.1007/978-3-031-16440-8_53 DOI: https://doi.org/10.1007/978-3-031-16440-8_53
9. Khalifa, M. and Albadawy, M. (2024). AI in diagnostic imaging: Revolutionizing accuracy and efficiency. Computer Methods and Programs in Biomedicine Update, 5(100146), pp.100146–100146. doi: https://doi.org/10.1016/j.cmpbup.2024.100146. DOI: https://doi.org/10.1016/j.cmpbup.2024.100146
10. Kothinti, R. R. (2024). The role of deep learning in radiology and medical imaging: improving diagnostic accuracy. International Journal of Academic Research and Development, 9, 128-143. https://www.academia.edu/download/121565934/IJNRD2409420_Paper_2024.pdf
11. Singh, S., Kumar, M., Kumar, A., Verma, B. K., Abhishek, K., & Selvarajan, S. (2024). Efficient pneumonia detection using Vision Transformers on chest X-rays. Scientific Reports, 14, 2487. DOI: https://doi.org/10.1038/s41598-024-52703-2
12. Mienye, I. D., Swart, T. G., Obaido, G., Jordan, M., & Ilono, P. (2025). Deep convolutional neural networks in medical image analysis: A review. Information, 16(3), 195. https://doi.org/10.3390/info16030195 DOI: https://doi.org/10.3390/info16030195
13. Natarajan, K., Muthusamy, S., Sha, M. S., Sadasivuni, K. K., Sekaran, S., Charles Gnanakkan, C. A. R., & A. Elngar, A. (2024). A novel method for the detection and classification of multiple diseases using transfer learning-based deep learning techniques with improved performance. Neural Computing and Applications, 36(30), 18979-18997. https://doi.org/10.1007/s00521-024-09900-x DOI: https://doi.org/10.1007/s00521-024-09900-x
14. Öksüz, C., Urhan, O., & Güllü, M. K. (2024). An integrated convolutional neural network with attention guidance for improved performance of medical image classification. Neural Computing and Applications, 36(4), 2067-2099. https://doi.org/10.1007/s00521-023-09164-x DOI: https://doi.org/10.1007/s00521-023-09164-x
15. Piffer, S., Ubaldi, L., Tangaro, S., Retico, A., & Talamonti, C. (2024). Tackling the small data problem in medical image classification with artificial intelligence: a systematic review. Progress in Biomedical Engineering, 6(3), 032001. https://doi.org/10.1088/2516-1091/ad525b DOI: https://doi.org/10.1088/2516-1091/ad525b
16. Rane, N. (2023). Transformers for medical image analysis: Applications, challenges, and future scope. Challenges and Future Scope (November 2, 2023). https://dx.doi.org/10.2139/ssrn.4622241 DOI: https://doi.org/10.2139/ssrn.4622241
17. Rayan, A. M., Adam, A., Al-Arabi, G., & Ahmed, M. R. (2025). The applications of X-ray technology in medical imaging: advances, challenges, and future perspectives (A review). Journal of Sustainable Food, Water, Energy and Environment, 1(2), 39- 61. https://journals.ekb.eg/article_454848_13d4a86fdddbc57bd9db3b245a746e11.pdf DOI: https://doi.org/10.21608/jsfw.2025.409882.1003
18. Salehi, A.W., Khan, S., Gupta, G., Alabduallah, B.I., Almjally, A., Alsolai, H., Siddiqui, T. and Mellit, A., 2023. A study of CNN and transfer learning in medical imaging: Advantages, challenges, future scope. Sustainability, 15(7), p.5930. https://doi.org/10.3390/su15075930 DOI: https://doi.org/10.3390/su15075930
19. Sallam, M., & Shnan, M. A. (2025). Enhancing semantic image retrieval using self-supervised learning: A label-efficient approach. Babylonian Journal of Machine Learning, 2025, 42- 60. https://doi.org/10.58496/BJML/2025/004 DOI: https://doi.org/10.58496/BJML/2025/004
20. Spinnato, P., Patel, D. B., Di Carlo, M., Bartoloni, A., Cevolani, L., Matcuk, G. R., & Crombé, A. (2022). Imaging of musculoskeletal soft-tissue infections in clinical practice: a comprehensive updated review. Microorganisms, 10(12), 2329. https://doi.org/10.3390/microorganisms10122329 DOI: https://doi.org/10.3390/microorganisms10122329
21. Tang, Y., Yang, D., Li, W., Roth, H., Landman, B., Xu, D., Nath, V., & Hatamizadeh, A. (2021). Self-Supervised Pre-Training of Swin Transformers for 3D Medical Image Analysis (Version 2). arXiv. https://doi.org/10.48550/ARXIV.2111.14791 DOI: https://doi.org/10.1109/CVPR52688.2022.02007
22. Wang, Y., Deng, Y., Zheng, Y., Chattopadhyay, P., & Wang, L. (2025). Vision transformers for image classification: A comparative survey. Technologies, 13(1), 32. https://doi.org/10.3390/technologies13010032 DOI: https://doi.org/10.3390/technologies13010032
23. Zhang, J., Li, F., Zhang, X., Wang, H., & Hei, X. (2024). Automatic Medical Image Segmentation with Vision Transformer. Applied Sciences, 14(7), 2741. https://doi.org/10.3390/app14072741 DOI: https://doi.org/10.3390/app14072741