SELF-SUPERVISED VISION TRANSFORMERS FOR EARLY DISEASE DETECTION IN LOW-RESOURCE MEDICAL IMAGING ENVIRONMENTS FRAMEWORK

Main Article Content

Anindya Bose
Denesh Sooriamoorthy
Kumaresu Murugasu

Abstract

This study proposes a Self-Supervised Vision Transformer (ViT) framework for early disease detection in low-resource medical imaging environments. The research addresses major challenges in healthcare systems, including limited labeled medical imaging datasets, inadequate computational resources, and a shortage of medical experts. Traditional Convolutional Neural Network (CNN)-based models generally require large annotated datasets and extensive computational power, limiting their applicability in resource-constrained healthcare settings. To overcome these limitations, the proposed framework integrates self-supervised learning techniques with a Vision Transformer architecture to improve medical image classification performance using limited labeled data. The framework uses chest X-ray images for pneumonia detection and employs preprocessing methods, including normalization, augmentation, resizing, and denoising, to improve image quality and consistency. Self-supervised learning techniques, including contrastive learning and masked image modeling, are used to learn meaningful image representations from unlabeled data. The Vision Transformer architecture captures both local and global contextual information through self-attention mechanisms, enhancing feature extraction and disease representation. Experimental evaluation is conducted using metrics such as accuracy, precision, recall, F1-score, confusion matrix, and ROC-AUC, with a comparative analysis against the ResNet50 CNN model. The results demonstrate that the proposed Vision Transformer framework achieves strong disease-sensitive classification performance and excellent recall for pneumonia detection, while reducing dependency on large annotated datasets. Although the CNN model achieved higher overall classification accuracy, the Vision Transformer demonstrated superior contextual learning and disease-discrimination capabilities. The study highlights the potential of self-supervised Vision Transformers for scalable, efficient, and cost-effective AI-assisted medical image analysis in low-resource healthcare environments.

Downloads

Download data is not yet available.

Article Details

Section

Articles

How to Cite

Bose, A., Sooriamoorthy, D., & Murugasu, K. (2025). SELF-SUPERVISED VISION TRANSFORMERS FOR EARLY DISEASE DETECTION IN LOW-RESOURCE MEDICAL IMAGING ENVIRONMENTS FRAMEWORK. Qubahan Journal of Medical Sciences, 2(1), 1-16. https://doi.org/10.48161/qjms.v2a84

References

1. Abdulrazzaq, M. M., Ramaha, N. T., Hameed, A. A., Salman, M., Yon, D. K., Fitriyani, N. L., ... & Lee, S. W. (2024). Consequential advancements of self-supervised learning (SSL) in deep learning contexts. Mathematics, 12(5), 758. https://doi.org/10.3390/math12050758 DOI: https://doi.org/10.3390/math12050758

2. Aburass, S., Dorgham, O., Al Shaqsi, J., Abu Rumman, M., & Al-Kadi, O. (2025). Vision Transformers in Medical Imaging: A Comprehensive Review of Advancements and Applications Across Multiple Diseases. Journal of Imaging Informatics in Medicine, 38(6), 3928–3971. https://doi.org/10.1007/s10278-025-01481-y DOI: https://doi.org/10.1007/s10278-025-01481-y

3. Chen, X., Xie, S., & He, K. (2021). An Empirical Study of Training Self-Supervised Vision Transformers (Version 4). arXiv. https://doi.org/10.48550/ARXIV.2104.02057 DOI: https://doi.org/10.1109/ICCV48922.2021.00950

4. Cleveland Clinic (2022). X-Ray: What It Is, Types, Preparation and Risks. [online] Cleveland Clinic. Available at: https://my.clevelandclinic.org/health/diagnostics/21818-x-ray.

5. Egbuna, I., Okei, N. C., Ajoku, E., Edobor, O., Issah, P., Olatokun, T., & Akinode, A. O. (2025). Advancing early disease detection with AI: innovations in medical imaging, EHR analytics, and wearable technologies. International Journal of Life Science Research Archive, 9, 120-140. https://doi.org/10.53771/ijlsra.2025.9.1.0049 DOI: https://doi.org/10.53771/ijlsra.2025.9.1.0049

6. Elyan, E., Vuttipittayamongkol, P., Johnston, P., Martin, K., McPherson, K., Moreno-García, C. F., ... & Sarker, M. M. K. (2022). Computer vision and machine learning for medical image analysis: recent advances, challenges, and way forward. Artificial Intelligence Surgery, 2(1), 24-45. http://dx.doi.org/10.20517/ais.2021.15 DOI: https://doi.org/10.20517/ais.2021.15

7. Huang, S.-C., Pareek, A., Jensen, M., Lungren, M. P., Yeung, S., & Chaudhari, A. S. (2023). Self-supervised learning for medical image classification: a systematic review and implementation guidelines. Npj Digital Medicine, 6(1). https://doi.org/10.1038/s41746-023-00811-0 DOI: https://doi.org/10.1038/s41746-023-00811-0

8. Jiang, J., Tyagi, N., Tringale, K., Crane, C., & Veeraraghavan, H. (2022). Self-supervised 3D Anatomy Segmentation Using Self-distilled Masked Image Transformer (SMIT). In Lecture Notes in Computer Science (pp. 556–566). Springer Nature Switzerland. https://doi.org/10.1007/978-3-031-16440-8_53 DOI: https://doi.org/10.1007/978-3-031-16440-8_53

9. Khalifa, M. and Albadawy, M. (2024). AI in diagnostic imaging: Revolutionizing accuracy and efficiency. Computer Methods and Programs in Biomedicine Update, 5(100146), pp.100146–100146. doi: https://doi.org/10.1016/j.cmpbup.2024.100146. DOI: https://doi.org/10.1016/j.cmpbup.2024.100146

10. Kothinti, R. R. (2024). The role of deep learning in radiology and medical imaging: improving diagnostic accuracy. International Journal of Academic Research and Development, 9, 128-143. https://www.academia.edu/download/121565934/IJNRD2409420_Paper_2024.pdf

11. Singh, S., Kumar, M., Kumar, A., Verma, B. K., Abhishek, K., & Selvarajan, S. (2024). Efficient pneumonia detection using Vision Transformers on chest X-rays. Scientific Reports, 14, 2487. DOI: https://doi.org/10.1038/s41598-024-52703-2

12. Mienye, I. D., Swart, T. G., Obaido, G., Jordan, M., & Ilono, P. (2025). Deep convolutional neural networks in medical image analysis: A review. Information, 16(3), 195. https://doi.org/10.3390/info16030195 DOI: https://doi.org/10.3390/info16030195

13. Natarajan, K., Muthusamy, S., Sha, M. S., Sadasivuni, K. K., Sekaran, S., Charles Gnanakkan, C. A. R., & A. Elngar, A. (2024). A novel method for the detection and classification of multiple diseases using transfer learning-based deep learning techniques with improved performance. Neural Computing and Applications, 36(30), 18979-18997. https://doi.org/10.1007/s00521-024-09900-x DOI: https://doi.org/10.1007/s00521-024-09900-x

14. Öksüz, C., Urhan, O., & Güllü, M. K. (2024). An integrated convolutional neural network with attention guidance for improved performance of medical image classification. Neural Computing and Applications, 36(4), 2067-2099. https://doi.org/10.1007/s00521-023-09164-x DOI: https://doi.org/10.1007/s00521-023-09164-x

15. Piffer, S., Ubaldi, L., Tangaro, S., Retico, A., & Talamonti, C. (2024). Tackling the small data problem in medical image classification with artificial intelligence: a systematic review. Progress in Biomedical Engineering, 6(3), 032001. https://doi.org/10.1088/2516-1091/ad525b DOI: https://doi.org/10.1088/2516-1091/ad525b

16. Rane, N. (2023). Transformers for medical image analysis: Applications, challenges, and future scope. Challenges and Future Scope (November 2, 2023). https://dx.doi.org/10.2139/ssrn.4622241 DOI: https://doi.org/10.2139/ssrn.4622241

17. Rayan, A. M., Adam, A., Al-Arabi, G., & Ahmed, M. R. (2025). The applications of X-ray technology in medical imaging: advances, challenges, and future perspectives (A review). Journal of Sustainable Food, Water, Energy and Environment, 1(2), 39- 61. https://journals.ekb.eg/article_454848_13d4a86fdddbc57bd9db3b245a746e11.pdf DOI: https://doi.org/10.21608/jsfw.2025.409882.1003

18. Salehi, A.W., Khan, S., Gupta, G., Alabduallah, B.I., Almjally, A., Alsolai, H., Siddiqui, T. and Mellit, A., 2023. A study of CNN and transfer learning in medical imaging: Advantages, challenges, future scope. Sustainability, 15(7), p.5930. https://doi.org/10.3390/su15075930 DOI: https://doi.org/10.3390/su15075930

19. Sallam, M., & Shnan, M. A. (2025). Enhancing semantic image retrieval using self-supervised learning: A label-efficient approach. Babylonian Journal of Machine Learning, 2025, 42- 60. https://doi.org/10.58496/BJML/2025/004 DOI: https://doi.org/10.58496/BJML/2025/004

20. Spinnato, P., Patel, D. B., Di Carlo, M., Bartoloni, A., Cevolani, L., Matcuk, G. R., & Crombé, A. (2022). Imaging of musculoskeletal soft-tissue infections in clinical practice: a comprehensive updated review. Microorganisms, 10(12), 2329. https://doi.org/10.3390/microorganisms10122329 DOI: https://doi.org/10.3390/microorganisms10122329

21. Tang, Y., Yang, D., Li, W., Roth, H., Landman, B., Xu, D., Nath, V., & Hatamizadeh, A. (2021). Self-Supervised Pre-Training of Swin Transformers for 3D Medical Image Analysis (Version 2). arXiv. https://doi.org/10.48550/ARXIV.2111.14791 DOI: https://doi.org/10.1109/CVPR52688.2022.02007

22. Wang, Y., Deng, Y., Zheng, Y., Chattopadhyay, P., & Wang, L. (2025). Vision transformers for image classification: A comparative survey. Technologies, 13(1), 32. https://doi.org/10.3390/technologies13010032 DOI: https://doi.org/10.3390/technologies13010032

23. Zhang, J., Li, F., Zhang, X., Wang, H., & Hei, X. (2024). Automatic Medical Image Segmentation with Vision Transformer. Applied Sciences, 14(7), 2741. https://doi.org/10.3390/app14072741 DOI: https://doi.org/10.3390/app14072741

Similar Articles

You may also start an advanced similarity search for this article.