EdgeFormerNet++: A Hybrid Attention-Guided Transformer Network for Accurate Retinal Layer Segmentation in OCT Images

Authors

  • Anju Thomas Department of Electronics and Communication Engineering, National Institute of Tiruchirappalli, Tiruchirappalli, Tamil Nadu https://orcid.org/0000-0002-2178-0531
  • Farhin Janath S. J. Department of Electronics and Electronics Engineering, Government Engineering College Barton Hill, Trivandrum, Kerala https://orcid.org/0009-0002-6426-7794
  • Vijayalakshmi P Department of Electronics and Electronics Engineering, Government Engineering College Barton Hill, Trivandrum, Kerala https://orcid.org/0009-0009-5451-5572
  • Nisan Pranavah Raja Department of Electronics and Communication Engineering, National Institute of Tiruchirappalli, Tiruchirappalli, Tamil Nadu https://orcid.org/0009-0006-4518-6787
  • Palanisamy P Department of Electronics and Communication Engineering, National Institute of Tiruchirappalli, Tiruchirappalli, Tamil Nadu https://orcid.org/0000-0003-3687-5944
  • Varun P Gopi Department of Electronics and Communication Engineering, National Institute of Tiruchirappalli, Tiruchirappalli, Tamil Nadu

Keywords:

Deep Learning, Edge Attention, Hybrid CNN-Transformer Architecture, Multi-class Semantic Segmentation, Optical Coherence Tomography (OCT), Res2Net, Retinal Layer Segmentation.

Abstract

Accurate retinal layer segmentation in OCT is crucial for the early detection and monitoring of retinal and neurodegenerative diseases, yet it remains challenging due to the thin, low-contrast structures and closely packed layers in multi-class settings. We propose EdgeFormerNet++, a hybrid framework that integrates Res2Net-based multi-scale encoding, lightweight Transformer bottlenecks for global context, and dual attention (CBAM and scSE) for spatial-channel recalibration, complemented by an edge attention module to refine layer boundaries. The model is evaluated on two public datasets (NR206 and MGU) covering macular and peripapillary regions, achieving state-of-the-art Dice scores of 92.34% and 83.28%, respectively. Qualitative and quantitative results demonstrate accurate delineation of complex retinal structures, supporting the use of EdgeFormerNet++ for automated OCT analysis and computer-assisted diagnosis.

Downloads

Download data is not yet available.

Author Biographies

Anju Thomas, Department of Electronics and Communication Engineering, National Institute of Tiruchirappalli, Tiruchirappalli, Tamil Nadu

Anju Thomas received the B.Tech. degree in Electronics and Communication Engineering from Vimal Jyothi Engineering College, Kannur, India, the M.Tech. degree in Signal Processing and Communication Engineering from Government Engineering College, Wayanada, India, and the Ph.D. degree in Electronics and Communication Engineering from the National Institute of Technology Tiruchirappalli, India. She is currently a Postdoctoral Fellow in the Department of Electronics and Communication Engineering at the National Institute of Technology Tiruchirappalli, India, under the DST-WISE Women Scientist Scheme. Her research interests include artificial intelligence, signal processing, medical image processing, deep learning, computer vision, and healthcare-oriented intelligent systems.

Farhin Janath S. J., Department of Electronics and Electronics Engineering, Government Engineering College Barton Hill, Trivandrum, Kerala

Farhin Janath S. J is currently pursuing the B.Tech. degree in Electronics and Communication Engineering at the Government Engineering College, Barton Hill, Trivandrum, Kerala, India. Her research interests include artificial intelligence and biomedical image processing.

Vijayalakshmi P, Department of Electronics and Electronics Engineering, Government Engineering College Barton Hill, Trivandrum, Kerala

Vijayalakshmi P is currently pursuing the B.Tech. degree in Electronics and Communication Engineering at the Government Engineering College, Barton Hill, Trivandrum, Kerala, India. Her research interests include artificial intelligence and biomedical image processing.

Nisan Pranavah Raja, Department of Electronics and Communication Engineering, National Institute of Tiruchirappalli, Tiruchirappalli, Tamil Nadu

Nisan Pranavah Raja received the B.E. degree in Electronics and Communication Engineering from Francis Xavier Engineering College, Tirunelveli, India, the M.E. degree in Communication Systems from Government College of Engineering, Tirunelveli, India, and currently received Ph.D. degree in Electronics and Communication Engineering from the National Institute of Technology Tiruchirappalli, India. His research interests include artificial intelligence, biomedical image processing, deep learning, and medical image analysis.

Palanisamy P, Department of Electronics and Communication Engineering, National Institute of Tiruchirappalli, Tiruchirappalli, Tamil Nadu

P. Palanisamy received the B.E. degree in Electronics and Communication Engineering from Bharathiar University, Coimbatore, India, the M.E. degree in Communication Systems, and the Ph.D. degree in Electronics and Communication Engineering from the National Institute of Technology Tiruchirappalli, India. He is currently a Professor in the Department of Electronics and Communication Engineering at the National Institute of Technology Tiruchirappalli, India. His research interests include signal processing, medical image analysis, machine learning, deep learning, computer vision, and artificial intelligence for healthcare applications.

Varun P Gopi, Department of Electronics and Communication Engineering, National Institute of Tiruchirappalli, Tiruchirappalli, Tamil Nadu

Varun P. Gopi received the B.Tech. degree in Electronics and Communication Engineering from Amal Jyothi College of Engineering, Kanjirappally, India, in 2007, and the M.Tech. degree from the College of Engineering Trivandrum, Kerala, India, in 2009. He obtained the Ph.D. degree in Electronics and Communication Engineering from the National Institute of Technology, Tiruchirappalli, India, in 2014. He is currently an Associate Professor in the Department of Electronics and Communication Engineering at the National Institute of Technology Tiruchirappalli, India. He is a Senior Member of IEEE. His research interests include artificial intelligence, machine learning, signal and image processing, medical image analysis, biomedical signal processing, and edge-AI systems for healthcare applications.

References

I. La´ıns, J. C. Wang, et al., “Retinal applications of swept source optical coherence tomography (oct) and optical coherence tomography angiography (octa),” Progress in Retinal and Eye Research, vol. 84, p. 100 951, 2021. DOI: 10 . 1016 / j . preteyeres . 2021 .100951.

J. Nathans, “Seeing is believing: The development of optical coherence tomography,” Proceedings of the National Academy of Sciences, vol. 120, no. 39, e2311129120, 2023. DOI:10.1073/pnas.2311129120.

A. London, I. Benhar, and M. Schwartz, “The retina as a window to the brain—from eye research to cns disorders,” Nature Reviews Neurology, vol. 9, no. 1, pp. 44–53, 2013. DOI: 10.1038/nrneurol.2012.227.

C. Rascun`a, A. Russo, et al., “Retinal thickness and microvascular pattern in early parkinson’s disease,” Frontiers in Neurology, vol. 11, p. 533 375, 2020. DOI: 10.3389/fneur.2020.533375.

H. Xie et al., “Arterial hypertension and retinal layer thickness: The beijing eye study,” British Journal of Ophthalmology, vol. 108, no. 1, pp. 105–111, 2024. DOI: 10.1136/bjo-2022-322229.

O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in International Conference on Medical Image Computing and Computer-Assisted Intervention, Springer, 2015, pp. 234–241. DOI: 10.1007/978-3-319-24574-4_28.

Z. Zhang, Q. Liu, and Y. Wang, “Road extraction by deep residual u-net,” IEEE Geoscience and Remote Sensing Letters, vol. 15, no. 5, pp. 749–753, 2018. DOI: 10.1109/LGRS.2018.2802944.

O. Oktay et al., “Attention u-net: Learning where to look for the pancreas,” in Medical Image Computing and Computer-Assisted Intervention – MICCAI 2018, vol. 11070, Springer, 2018, pp. 121–130. DOI: 10 .48550/arXiv.1804.03999.

H. Cao et al., “Swin-unet: Unet-like pure transformer for medical image segmentation,” in Computer Vision – ECCV 2022 Workshops, Springer, 2022, pp. 205–218.

DOI: 10.1007/978-3-03125066-8 9.

S. Jardim, J. Ant´onio, and C. Mora, “Image thresholding approaches for medical image segmentation — short literature review,” Procedia Computer Science, vol. 219, pp. 1485–1492, 2023. DOI: 10.1016/j.procs.2023.01.439.

S. Pare, A. Kumar, G. K. Singh, and V. Bajaj, “Image segmentation using multilevel thresholding: A research review,” Iranian Journal of Science and Technology, Transactions of Electrical Engineering, vol. 44, no. 1, pp. 1–29, 2020. DOI: 10.1007/s40998-019-00251-1.

J. Liu et al., “Automated retinal boundary segmentation of optical coherence tomography images using an improved canny operator,” Scientific Reports, vol. 12, no. 1, p. 1412, 2022. DOI: 10.1038/s41598-022-05550-y.

Y. Chen, P. Ge, G. Wang, G. Weng, and H. Chen, “An overview of intelligent image segmentation using active contour models,” Intelligent Robotics, vol. 3, no. 1, pp. 23–55, 2023. DOI: 10.20517/ir.2023.02.

K. Ramesh, G. K. Kumar, K. Swapna, D. Datta, and S. S. Rajest, “A review of medical image segmentation algorithms,” EAI Endorsed Transactions on Pervasive Health & Technology, vol. 7, no. 27, 2021. DOI: 10.4108/eai.12-4-2021.169184.

M. A. Chandra and S. S. Bedi, “Survey on svm and their application in image classification,” International Journal of Information Technology, vol. 13, no. 5, pp. 1– 11, 2021. DOI: 10.1007/s41870-017-0080-1.

T. Zhu, “Analysis on the applicability of the random forest,” in Journal of Physics: Conference Series, vol. 1607, IOP Publishing, 2020, p. 012 123. DOI: 10.1088/1742-6596/1607/1/012123.

R. Azad et al., “Medical image segmentation review: The success of u-net,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 12, pp. 10 076–10 095, 2024. DOI:10.1109/TPAMI.2024.3435571.

A. G. Roy et al., “Relaynet: Retinal layer and fluid segmentation of macular optical coherence tomography using fully convolutional networks,” Biomedical Optics Express, vol. 8, no. 8, pp. 3620–3637, 2017. DOI: 10.1364/BOE.8.003627.

B. M. Zipporah and D. F. X. Christopher, “Residual u-net architecture for retinal layers in oct images with choroidal neovascularization,” International Journal of Electrical and Electronics Engineering, vol. 10, no. 5, pp. 205–212, 2023. DOI: 10.14445/23488379/IJEEEV10I5P119

B. Fazekas et al., “Sd-layernet: Semi-supervised retinal layer segmentation in oct using disentangled representation with anatomical priors,” in Medical Image Computing and Computer-Assisted Intervention – MICCAI 2022 Springer, 2022, pp. 320–329. DOI: 10.1007/978-3-031-16452-1_31.

J. Chen et al., “Transunet: Rethinking the u-net architecture design for medical image segmentation through the lens of transformers,” Medical Image Analysis, vol. 97, p. 103 280, 2024. DOI: 10.1016/j.media.2024.103280.

Y. Gao, M. Zhou, and D. N. Metaxas, “Utnet: A hybrid transformer architecture for medical image segmentation,” in Medical Image Computing and Computer-Assisted Intervention – MICCAI 2021,

Springer, 2021, pp. 61–71. DOI: 10.1007/978-3-030-87199-4 6.

A. Hatamizadeh et al., “Unetr: Transformers for 3d medical image segmentation,” in Proceedings of the IEEE/CVF winter conference on applications of computer vision, 2022, pp. 574–584. DOI: https://doi.org/10.48550/arXiv.2103.10504.

J. Li et al., “Multi-scale gcn-assisted two-stage network for joint segmentation of retinal layers and discs in peripapillary oct images,” Biomedical Optics Express, vol. 12, no. 4, pp. 2204–2220, 2021. DOI: 10 . 1364 /BOE.417212.

E. Liu et al., “Mt net: A multi-scale framework using the transformer block for retina layer segmentation,”Photonics, vol. 11, p. 607, 2024. DOI: 10 . 3390 /photonics11070607.

X. He et al., “Lightweight retinal layer segmentation with global reasoning,” IEEE transactions on instrumentation and measurement, vol. 73, pp. 1–14, 2024. DOI: 10.1109/TIM.2024.3400305.

J. Hao, H. Li, S. Lu, Z. Li, and W. Zhang, “General retinal layer segmentation in oct images via reinforcement constraint,” Computerized Medical Imaging and Graphics, vol. 120, p. 102 480, 2025. DOI: 10.1016/j.compmedimag.2024.102480.

H. Peng, Y. Yu, and S. Yu, “Re-thinking the effectiveness of batch normalization and beyond,” IEEE transactions on pattern analysis and machine intelligence, vol. 46, no. 1, pp. 465–478, 2023. DOI: 10.1109/TPAMI.2023.3319005.

K. Hara, D. Saito, and H. Shouno, “Analysis of function of rectified linear unit used in deep learning,” in 2015 international joint conference on neural networks (IJCNN), IEEE, 2015, pp. 1–8. DOI: 10.1109/IJCNN.2015.7280578.

S.-H. Gao, M.-M. Cheng, K. Zhao, X.-Y. Zhang, M.-H. Yang, and P. Torr, “Res2net: A new multi-scale backbone architecture,” IEEE transactions on pattern analysis and machine intelligence, vol. 43, pp. 652–662, 2019. DOI: 10.1109/TPAMI.2019.2938758.

S. Woo, J. Park, J.-Y. Lee, and I. S. Kweon, “Cbam: Convolutional block attention module,” in Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 3–19. DOI: 10.1007/978- 3- 030-01234-2_1.

A. G. Roy, N. Navab, and C. Wachinger, “Recalibrating fully convolutional networks with spatial and channel ‘squeeze & excitation’ blocks,” in Medical Image Computing and Computer-Assisted Intervention – MICCAI 2018, vol. 11071, Springer, 2018, pp. 733–741. DOI:10.1109/TMI.2018.2867261.

P. Gholami, P. Roy, and V. Lakshminarayanan, “Octid: Optical coherence tomography image database,” Data in Brief, vol. 28, p. 104 882, 2020. DOI: 10 . 1016 / j .compeleceng.2019.106532.

A. Buslaev, V. I. Iglovikov, E. Khvedchenya, A. Parinov, A. Druzhinin, and A. A. Kalinin, “Albumentations: Fast and flexible image augmentations,” Information, vol. 11, no. 2, p. 125, 2020. DOI: 10 . 3390 /info11020125.

F. Milletari, N. Navab, and S.-A. Ahmadi, “V-net: Fully convolutional neural networks for volumetric medical image segmentation,” in Proceedings of the Fourth International Conference on 3D Vision (3DV), IEEE, 2016, pp. 565–571. DOI: 10.1109/3DV.2016.79.

O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in Medical Image Computing and Computer- Assisted Intervention (MICCAI), vol. 9351, Springer, 2015, pp. 234–241. DOI: 10.1007/978-3-319-24574-4_28.

F. Isensee, P. F. Jaeger, S. A. A. Kohl, J. Petersen, and K. H. Maier-Hein, “Nnu-net: A self-configuring method for deep learning-based biomedical image segmentation,” Nature Methods, vol. 18, no. 2, pp. 203–211, 2021. DOI: 10.1038/s41592-020-01008-z.

X. He et al., “Exploiting multi-granularity visual features for retinal layer segmentation in human eyes,” Frontiers in Bioengineering and Biotechnology, vol. 11, p. 1 191 803, 2023. DOI: 10.3389/fbioe.2023.1191803.

Published

2026-08-30

How to Cite

Thomas, A., S. J., F. J. ., P, . V., Raja, N. P., P, P., & P Gopi, V. (2026). EdgeFormerNet++: A Hybrid Attention-Guided Transformer Network for Accurate Retinal Layer Segmentation in OCT Images. IEEE Latin America Transactions, 24(10), 1093–1104. Retrieved from https://latamt.ieeer9.org/index.php/transactions/article/view/10551