Preview

Informatics

Advanced search

Modified DeepLabV3+ architecture with global context integration and supervised contrastive learning for semantic segmentation of ultra-high-resolution image

https://doi.org/10.37661/1/1816-0301-2026-23-3-7-23

Abstract

Objectives. We propose a modern method for semantic segmentation of ultra-high-resolution (4K) video frames in the field of Earth remote sensing using a modification of the DeepLabV3+ convolutional neural network.

Methods. The method is aimed at solving two critical problems: the limited receptive field of the model when working with local image fragments due to GPU memory constraints, and the poor separation of classes with similar color characteristics and texture. The baseline architecture was supplemented with two attention mechanism modules to improve object localization, and the standard cross-entropy loss function was supplemented with a Supervised Contrastive Learning (SupCon) algorithm for better separation of similar classes. Feature-wise Linear Modulation (FiLM) was embedded into the Atrous Spatial Pyramid Pooling (ASPP) module to introduce global context in the form of features extracted from the original image.

Results. The results of the experiments showed that the proposed approach successfully minimizes artifacts at class boundaries, has a high ability to distinguish visually similar classes, and improves segmentation accuracy under the conditions of a specific dataset. The proposed architecture outperforms the baseline by 0,45 % in terms of mean Intersection over Union.

Conclusion. The developed modification of DeepLabV3+ effectively addresses the challenges of semantic segmentation of ultra-high-resolution images in the field of Earth remote sensing, achieving high accuracy with a negligible increase in computational complexity.

About the Authors

A. A. Kozlov
Belarusian State University
Belarus

Anatoly A. Kozlov , Postgraduate Student at the Department of Web Technologies and Computer Modeling

av. Nezavisimosti, 4, Minsk, 220030



S. V. Ablameyko
Belarusian State University
Belarus

Sergey V. Ablameyko , Dr. Sci. (Eng.), Acad. of the National Academy of Sciences of Belarus, Prof. at the Department of Web Technologies and Computer Modeling

av. Nezavisimosti, 4, Minsk, 220030



References

1. Jiang C., Chervan A. N. Identification of land use structure in the zone of influence of the Soligorsk potash plant based on remote sensing data. Prirodopolzovanie [Nature Management], 2025, no. 1, рр. 51–63 (In Russ.).

2. Roche C., Thygesen K., Baker E. Mine tailings storage: Safety is no accident. Rapid Response Assessment. United Nations Environment Programme, GRID-Arendal. Arendal, 2017, рр. 6–10.

3. Rotta L. H. S., Alcântara E., Park E., Negri R. G., Lin Y. N. The 2019 Brumadinho tailings dam collapse: Possible cause and impacts of the worst human and environmental disaster in Brazil. International Journal of Applied Earth Observation and Geoinformation, 2020, vol. 90, art. 102119. Available at: https://www.sciencedirect.com/science/article/pii/S0303243420300192?via%3Dihub (accessed: 16.03.2026). https://doi.org/10.1016/j.jag.2020.102119.

4. Santamarina J. C., Torres-Cruz L. A., Bachus R. C. Why coal ash and tailings dam disasters occur. Science, 2019, vol. 364, iss. 6440, рр. 526–528. https://doi.org/10.1126/science.aax1927.

5. Balaniuk R., Isupova O., Reece S. Mining and tailings dam detection in satellite imagery using deep learning. Sensors, 2020, vol. 20, no. 23, art. 6936. Available at: https://www.mdpi.com/1424-8220/20/23/6936 (accessed 16.03.2026). https://doi.org/10.3390/s20236936.

6. Stelmakh V. I. Environmental problems of potash salt mining and some solutions. Obespechenie jekonomicheskoj bezopasnosti Respubliki Belarus': teorija i praktika: materialy Respublikanskoj nauchno-prakticheskoj konferencii, Minsk, 19 dekabrja 2019 g. [Ensuring the Economic Security of the Republic of Belarus: Theory and Practice: Proceedings of the Republican Scientific-practical Conference, Minsk, 19 December 2019]. Minsk, Akademija Ministerstva vnutrennih del Respubliki Belarus', 2019, рр. 103–105 (In Russ.).

7. Long J., Shelhamer E., Darrell T. Fully convolutional networks for semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, Massachusetts, USA, 7–12 June 2015. Boston, 2015, рр. 3431–3440. https://doi.org/10.1109/CVPR.2015.7298965.

8. Ronneberger O., Fischer P., Brox T. U-Net: Convolutional networks for biomedical image segmentation. Medical Image Computing and Computer-Assisted Intervention – MICCAI 2015 : 18th International Conference, Munich, 5–9 October 2015. In N. Navab, J. Hornegger, W. M. Wells, A. F. Frangi (eds.). Cham, 2015, pt. III, рр. 234–241 (Lecture Notes in Computer Science; vol. 9351). https://doi.org/10.1007/978-3-319-24574-4_28.

9. H. Zhao, Shi J., Qi X., Wang X., Jia J. Pyramid scene parsing network. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, Hawaii, USA, 21–26 July 2017. Honolulu, 2017, рр. 2881–2890. https://doi.org/10.1109/CVPR.2017.660.

10. Chen L.-C., Papandreou G., Kokkinos I., Murphy K., Yuille A. L. Semantic image segmentation with deep convolutional nets and fully connected CRFs. International Conference on Learning Representations (ICLR), San Diego, CA, USA, 7–9 May 2015. San Diego, 2015, рр. 1–14. https://doi.org/10.48550/arXiv.1412.7062.

11. Chen L.-C., Papandreou G., Kokkinos I., Murphy K., Yuille A. L. DeepLab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected CRFs. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2018, vol. 40, iss. 4, рр. 834–848. https://doi.org/10.1109/TPAMI.2017.2699184.

12. Chen L.-C., Papandreou G., Schroff F., Adam H. Rethinking atrous convolution for semantic image segmentation, 2017. Available at: https://arxiv.org/pdf/1706.05587 (accessed: 16.03.2026). https://doi.org/10.48550/arXiv.1706.05587.

13. Chen L.-C., Zhu Y., Papandreou G., Schroff F., Adam H. Encoder-decoder with atrous separable convolution for semantic image segmentation. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018. Munich, 2018, рр. 801–818. https://doi.org/10.1007/978-3-030-01234-2_49.

14. Dosovitskiy A., Beyer L., Kolesnikov A., Weissenborn D., Zhai X., …, Houlsby N. An image is worth 16x16 words: Transformers for image recognition at scale. International Conference on Learning Representations (ICLR), Vienna, Austria, 4 May 2021. Vienna, 2021, рр. 1–21. https://doi.org/10.48550/arXiv.2010.11929.

15. Zheng S., Lu J., Zhao H., Zhu X., Luo Z., Wang Y. Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 20–25 June 2021. Nashville, 2021, рр. 6881–6890. https://doi.org/10.1109/CVPR46437.2021.00681.

16. Godase V. V., Takale S., Ghodake R., Mulani A. Attention mechanisms in semantic segmentation of remote sensing images. Journal of Advancement in Electronics Signal Processing, 2025, vol. 2, рр. 45–58.

17. Hu J., Shen L., Sun G. Squeeze-and-excitation networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–22 June 2018. Salt Lake City, 2018, рр. 7132–7141. https://doi.org/10.1109/CVPR.2018.00745.

18. Woo S., Park J., Lee J. Y., Kweon I. S. CBAM: Convolutional block attention module. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018. Munich, 2018, рр. 3–19. https://doi.org/10.1007/978-3-030-01234-2_1.

19. Hou Q., Zhou D., Feng J. Coordinate attention for efficient mobile network design. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 20–25 June 2021. Nashville, 2021, рр. 13713–13722. https://doi.org/10.1109/CVPR46437.2021.01350.

20. Chen T., Kornblith S., Swersky K., Norouzi M., Hinton G. Big self-supervised models are strong semi-supervised learners. Advances in Neural Information Processing Systems (NeurIPS), 2020, vol. 33, рр. 20735–20747. https://doi.org/10.48550/arXiv.2006.10029.

21. Hu H., Wang X., Zhang Y., Chen Q., Guan Q. A comprehensive survey on contrastive learning. Neurocomputing, 2024, vol. 610, art. 128645. Available at: https://www.sciencedirect.com/science/article/abs/pii/S0925231224014164?via%3Dihub (accessed 16.03.2026). https://doi.org/10.1016/j.neucom.2024.128645.

22. Khosla P., Teterwak P., Wang C., Sarna A., Tian Y., …, Krishnan D. Supervised contrastive learning. Advances in Neural Information Processing Systems (NeurIPS), 2020, vol. 33, рр. 18661–18673. https://doi.org/10.48550/arXiv.2004.11362.

23. He K., Zhang X., Ren S., Sun J. Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27–30 June 2016. Las Vegas, 2016, рр. 770–778. https://doi.org/10.1109/CVPR.2016.90.

24. Bui P.-N., Le D.-T., Bum J., Kim S., Song S. J., Choo H. Comparative analysis of deep learning architectures for medical image classification. Bioengineering, 2023, vol. 10, iss. 10, art. 1249. Available at: https://pmc.ncbi.nlm.nih.gov/articles/PMC10669434/ (accessed 16.03.2026). https://doi.org/10.3390/bioengineering10101249.

25. Kansal K., Singh A., Srivastava D., Kansal K. ResNet-based deep learning approaches for histopathological breast cancer classification. 2025 International Conference on Artificial Intelligence and Emerging Technologies (ICAIET), Bhubaneswar, Odisha, India, 28–30 August 2025. Bhubaneswar, 2025, рр. 1–4. https://doi.org/10.1109/ICAIET.2025.11210953.

26. Kordemir M., Cevik K. K., Bozkurt A. A mask R-CNN approach for detection and classification of brain tumours from MR images. Computer Methods in Biomechanics and Biomedical Engineering: Imaging & Visualization, 2024, vol. 11, iss. 7. Available at: https://www.tandfonline.com/doi/full/10.1080/21681163.2023.2301391 (accessed 16.03.2026). https://doi.org/10.1080/21681163.2023.2301391.

27. Zitouni M. S., Alkhatib M. Q., Aburaed N., Ahmad H. A comparative analysis of attention mechanisms in 3D CNN-based hyperspectral image super-resolution. 2024 14th Workshop on Hyperspectral Imaging and Signal Processing: Evolution in Remote Sensing (WHISPERS), Helsinki, Finland, 09–11 December 2024. Helsinki, 2024, рр. 1–5. https://doi.org/10.1109/WHISPERS.2024.10876507.

28. Yang Z., Wu Q., Zhang F., Zhang X., Chen X., Gao Y. A new semantic segmentation method for remote sensing images integrating coordinate attention and SPD-Conv. Symmetry, 2023, vol. 15, iss. 5, art. 1037. Available at: https://www.mdpi.com/2073-8994/15/5/1037 (accessed 16.03.2026). https://doi.org/10.3390/sym15051037.

29. Li Z. A comparative study of Contrastive Self-Supervised Learning (CSSL): Methods, technologies, and applications. Transactions on Computer Science and Intelligent Systems Research, 2025, vol. 10, рр. 92–97. https://doi.org/10.62051/2hgr2j26.

30. Wang Y., Albrecht C. M., Zhu X. X. Multi-label guided supervised contrastive learning for earth observation pretraining. 2024 IEEE International Geoscience and Remote Sensing Symposium (IGARSS), Athens, Greece, 7–12 July 2024. Athens, 2024, рр. 7568–7571. https://doi.org/10.1109/IGARSS53475.2024.10642289.

31. Howard A., Sandler M., Chen B., Wang W., Chen L.-C., Tan M. Searching for MobileNetV3. Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Korea (South), 27 October – 02 November 2019. Seoul, 2019, рр. 1314–1324. https://doi.org/10.1109/ICCV.2019.00140.

32. Perez E., Strub F., Vries H. de, Dumoulin V., Courville A. FiLM: Visual reasoning with a general conditioning layer. Proceedings of the AAAI Conference on Artificial Intelligence, New Orleans, Louisiana, USA, 2–7 February 2018. New Orleans, 2018, vol. 32, iss. 1, рр. 3942–3951. https://doi.org/10.1609/aaai.v32i1.11823.


Review

For citations:


Kozlov A.A., Ablameyko S.V. Modified DeepLabV3+ architecture with global context integration and supervised contrastive learning for semantic segmentation of ultra-high-resolution image. Informatics. 2026;23(3):7-23. (In Russ.) https://doi.org/10.37661/1/1816-0301-2026-23-3-7-23

Views: 13

JATS XML


Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 License.


ISSN 1816-0301 (Print)
ISSN 2617-6963 (Online)