<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.3 20210610//EN" "JATS-journalpublishing1-3.dtd">
<article article-type="research-article" dtd-version="1.3" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xml:lang="ru"><front><journal-meta><journal-id journal-id-type="publisher-id">inform</journal-id><journal-title-group><journal-title xml:lang="ru">Информатика</journal-title><trans-title-group xml:lang="en"><trans-title>Informatics</trans-title></trans-title-group></journal-title-group><issn pub-type="ppub">1816-0301</issn><issn pub-type="epub">2617-6963</issn><publisher><publisher-name>UIIP NASB</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.37661/1/1816-0301-2026-23-3-7-23</article-id><article-id custom-type="elpub" pub-id-type="custom">inform-1403</article-id><article-categories><subj-group subj-group-type="heading"><subject>Research Article</subject></subj-group><subj-group subj-group-type="section-heading" xml:lang="ru"><subject>ОБРАБОТКА СИГНАЛОВ, ИЗОБРАЖЕНИЙ, РЕЧИ, ТЕКСТА И РАСПОЗНАВАНИЕ ОБРАЗОВ</subject></subj-group><subj-group subj-group-type="section-heading" xml:lang="en"><subject>SIGNAL, IMAGE, SPEECH, TEXT PROCESSING AND PATTERN RECOGNITION</subject></subj-group></article-categories><title-group><article-title>Модифицированная архитектура DeepLabV3+ с интеграцией глобального контекста и контрастивного обучения с учителем для семантической сегментации изображений сверхвысокого разрешения</article-title><trans-title-group xml:lang="en"><trans-title>Modified DeepLabV3+ architecture with global context integration and supervised contrastive learning for semantic segmentation of ultra-high-resolution image</trans-title></trans-title-group></title-group><contrib-group><contrib contrib-type="author" corresp="yes"><name-alternatives><name name-style="eastern" xml:lang="ru"><surname>Козлов</surname><given-names>А. А.</given-names></name><name name-style="western" xml:lang="en"><surname>Kozlov</surname><given-names>A. A.</given-names></name></name-alternatives><bio xml:lang="ru"><p>Козлов Анатолий Анатольевич , аспирант кафедры веб-технологий и компьютерного моделирования</p><p>пр. Независимости, 4, Минск, 220030</p></bio><bio xml:lang="en"><p>Anatoly A. Kozlov , Postgraduate Student at the Department of Web Technologies and Computer Modeling</p><p>av. Nezavisimosti, 4, Minsk, 220030</p></bio><email xlink:type="simple">anatoly.kozlov.ak@yandex.ru</email><xref ref-type="aff" rid="aff-1"/></contrib><contrib contrib-type="author" corresp="yes"><name-alternatives><name name-style="eastern" xml:lang="ru"><surname>Абламейко</surname><given-names>С. В.</given-names></name><name name-style="western" xml:lang="en"><surname>Ablameyko</surname><given-names>S. V.</given-names></name></name-alternatives><bio xml:lang="ru"><p>Абламейко Сергей Владимирович , доктор технических наук, академик Национальной академии наук Республики Беларусь, профессор кафедры веб-технологий и компьютерного моделирования</p><p>пр. Независимости, 4, Минск, 220030</p></bio><bio xml:lang="en"><p>Sergey V. Ablameyko , Dr. Sci. (Eng.), Acad. of the National Academy of Sciences of Belarus, Prof. at the Department of Web Technologies and Computer Modeling</p><p>av. Nezavisimosti, 4, Minsk, 220030</p></bio><email xlink:type="simple">ablameyko@bsu.by</email><xref ref-type="aff" rid="aff-1"/></contrib></contrib-group><aff-alternatives id="aff-1"><aff xml:lang="ru"><institution>Белорусский государственный университет</institution><country>Беларусь</country></aff><aff xml:lang="en"><institution>Belarusian State University</institution><country>Belarus</country></aff></aff-alternatives><pub-date pub-type="collection"><year>2026</year></pub-date><pub-date pub-type="epub"><day>26</day><month>09</month><year>2026</year></pub-date><volume>23</volume><issue>3</issue><fpage>7</fpage><lpage>23</lpage><permissions><copyright-statement>Copyright &amp;#x00A9; Козлов А.А., Абламейко С.В., 2026</copyright-statement><copyright-year>2026</copyright-year><copyright-holder xml:lang="ru">Козлов А.А., Абламейко С.В.</copyright-holder><copyright-holder xml:lang="en">Kozlov A.A., Ablameyko S.V.</copyright-holder><license xml:lang="ru" license-type="creative-commons-attribution" xlink:href="https://creativecommons.org/licenses/by/4.0/" xlink:type="simple"><license-p>Данная работа распространяется под лицензией Creative Commons Attribution 4.0.</license-p></license><license xml:lang="en" license-type="creative-commons-attribution" xlink:href="https://creativecommons.org/licenses/by/4.0/" xlink:type="simple"><license-p>This work is licensed under a Creative Commons Attribution 4.0 License.</license-p></license></permissions><self-uri xlink:href="https://inf.grid.by/jour/article/view/1403">https://inf.grid.by/jour/article/view/1403</self-uri><abstract><sec><title>Цели</title><p>Цели. Целью исследования является разработка современного метода семантической сегментации видеокадров сверхвысокого разрешения (4К) в области задач дистанционного зондирования Земли (ДЗЗ) с помощью модификации сверточной нейронной сети DeepLabV3+.</p></sec><sec><title>Методы</title><p>Методы. Метод направлен на решение двух критических проблем: ограниченности рецептивного поля модели при работе с локальными фрагментами изображения в силу больших затрат видеопамяти графических процессоров и некачественного разделения классов со схожими цветовыми характеристиками и текстурой. Базовая архитектура была дополнена двумя модулями механизма внимания для улучшения локализации объектов, а стандартная функция потерь cross entropy алгоритмом контрастивного обучения с учителем (Supervised Contrastive Learning, SupCon) для лучшего разделения схожих классов. В Atrous Spatial Pyramid Pooling (ASPP)-модуль была встроена линейная модуляция (Feature-wise Linear Modulation, FiLM) для внедрения глобального контекста в виде признаков, полученных из исходного изображения.</p></sec><sec><title>Результаты</title><p>Результаты. Результаты проведенных экспериментов показали, что предложенный подход успешно справляется с минимизацией артефактов на границах классов, обладает высокой способностью различать визуально схожие классы, а также повышает точность сегментации в условиях специфического набора данных. Предложенная архитектура превосходит базовую по метрике mean Intersection over Union на 0,45 %.</p></sec><sec><title>Заключение</title><p>Заключение. Разработанная модификация DeepLabV3+ позволяет эффективно решать задачи семантической сегментации изображений сверхвысокого разрешения в области ДЗЗ, обеспечивая высокую точность при незначительном увеличении вычислительной сложности.</p></sec></abstract><trans-abstract xml:lang="en"><sec><title>Objectives</title><p>Objectives. We propose a modern method for semantic segmentation of ultra-high-resolution (4K) video frames in the field of Earth remote sensing using a modification of the DeepLabV3+ convolutional neural network.</p></sec><sec><title>Methods</title><p>Methods. The method is aimed at solving two critical problems: the limited receptive field of the model when working with local image fragments due to GPU memory constraints, and the poor separation of classes with similar color characteristics and texture. The baseline architecture was supplemented with two attention mechanism modules to improve object localization, and the standard cross-entropy loss function was supplemented with a Supervised Contrastive Learning (SupCon) algorithm for better separation of similar classes. Feature-wise Linear Modulation (FiLM) was embedded into the Atrous Spatial Pyramid Pooling (ASPP) module to introduce global context in the form of features extracted from the original image.</p></sec><sec><title>Results</title><p>Results. The results of the experiments showed that the proposed approach successfully minimizes artifacts at class boundaries, has a high ability to distinguish visually similar classes, and improves segmentation accuracy under the conditions of a specific dataset. The proposed architecture outperforms the baseline by 0,45 % in terms of mean Intersection over Union.</p></sec><sec><title>Conclusion</title><p>Conclusion. The developed modification of DeepLabV3+ effectively addresses the challenges of semantic segmentation of ultra-high-resolution images in the field of Earth remote sensing, achieving high accuracy with a negligible increase in computational complexity.</p></sec></trans-abstract><kwd-group xml:lang="ru"><kwd>дистанционное зондирование Земли</kwd><kwd>сверточные нейронные сети</kwd><kwd>семантическая сегментация</kwd><kwd>контрастивное обучение</kwd><kwd>мониторинг терриконов</kwd><kwd>механизм внимания</kwd></kwd-group><kwd-group xml:lang="en"><kwd>Earth remote sensing</kwd><kwd>convolutional neural networks</kwd><kwd>semantic segmentation</kwd><kwd>contrastive learning</kwd><kwd>tailings pond monitoring</kwd><kwd>attention mechanism</kwd></kwd-group></article-meta></front><back><ref-list><title>References</title><ref id="cit1"><label>1</label><citation-alternatives><mixed-citation xml:lang="ru">Цзян, Ч. Выявление структуры землепользования в зоне влияния Солигорского калийного комбината по данным дистанционного зондирования / Ч. Цзян, А. Н. Червань // Природопользование. – 2025. –№ 1. – С. 51–63.</mixed-citation><mixed-citation xml:lang="en">Jiang C., Chervan A. N. Identification of land use structure in the zone of influence of the Soligorsk potash plant based on remote sensing data. Prirodopolzovanie [Nature Management], 2025, no. 1, рр. 51–63 (In Russ.).</mixed-citation></citation-alternatives></ref><ref id="cit2"><label>2</label><citation-alternatives><mixed-citation xml:lang="ru">Roche, C. Mine tailings storage: Safety is no accident / C. Roche, K. Thygesen, E. Baker // Rapid Response Assessment / United Nations Environment Programme, GRID-Arendal. – Arendal, 2017. – P. 6–10.</mixed-citation><mixed-citation xml:lang="en">Roche C., Thygesen K., Baker E. Mine tailings storage: Safety is no accident. Rapid Response Assessment. United Nations Environment Programme, GRID-Arendal. Arendal, 2017, рр. 6–10.</mixed-citation></citation-alternatives></ref><ref id="cit3"><label>3</label><citation-alternatives><mixed-citation xml:lang="ru">The 2019 Brumadinho tailings dam collapse: Possible cause and impacts of the worst human and environmental disaster in Brazil / L. H. S. Rotta, E. Alcântara, E. Park [et al.] // International Journal of Applied Earth Observation and Geoinformation. – 2020. – Vol. 90. – Art. 102119. – URL: https://www.sciencedirect.com/science/article/pii/S0303243420300192?via%3Dihub (date of access: 16.03.2026). – https://doi.org/10.1016/j.jag.2020.102119.</mixed-citation><mixed-citation xml:lang="en">Rotta L. H. S., Alcântara E., Park E., Negri R. G., Lin Y. N. The 2019 Brumadinho tailings dam collapse: Possible cause and impacts of the worst human and environmental disaster in Brazil. International Journal of Applied Earth Observation and Geoinformation, 2020, vol. 90, art. 102119. Available at: https://www.sciencedirect.com/science/article/pii/S0303243420300192?via%3Dihub (accessed: 16.03.2026). https://doi.org/10.1016/j.jag.2020.102119.</mixed-citation></citation-alternatives></ref><ref id="cit4"><label>4</label><citation-alternatives><mixed-citation xml:lang="ru">Santamarina, J. C. Why coal ash and tailings dam disasters occur / J. C. Santamarina, L. A. Torres-Cruz, R. C. Bachus // Science. – 2019. – Vol. 364, iss. 6440. – P. 526–528. – https://doi.org/10.1126/science.aax1927.</mixed-citation><mixed-citation xml:lang="en">Santamarina J. C., Torres-Cruz L. A., Bachus R. C. Why coal ash and tailings dam disasters occur. Science, 2019, vol. 364, iss. 6440, рр. 526–528. https://doi.org/10.1126/science.aax1927.</mixed-citation></citation-alternatives></ref><ref id="cit5"><label>5</label><citation-alternatives><mixed-citation xml:lang="ru">Mining and tailings dam detection in satellite imagery using deep learning / R. Balaniuk, O. Isupova, S. Reece // Sensors. – 2020. – Vol. 20, no. 23. – Art. 6936. – URL: https://www.mdpi.com/1424-8220/20/23/6936 (date of access: 16.03.2026). – https://doi.org/10.3390/s20236936.</mixed-citation><mixed-citation xml:lang="en">Balaniuk R., Isupova O., Reece S. Mining and tailings dam detection in satellite imagery using deep learning. Sensors, 2020, vol. 20, no. 23, art. 6936. Available at: https://www.mdpi.com/1424-8220/20/23/6936 (accessed 16.03.2026). https://doi.org/10.3390/s20236936.</mixed-citation></citation-alternatives></ref><ref id="cit6"><label>6</label><citation-alternatives><mixed-citation xml:lang="ru">Стельмах, В. И. Экологические проблемы добычи калийных солей и некоторые пути их преодоления / В. И. Стельмах // Обеспечение экономической безопасности Республики Беларусь: теория и практика : материалы Респ. науч.-практ. конф., Минск, 19 дек. 2019 г. / Акад. М-ва внутр. дел Респ. Беларусь. – Мн., 2019. – С. 103–105.</mixed-citation><mixed-citation xml:lang="en">Stelmakh V. I. Environmental problems of potash salt mining and some solutions. Obespechenie jekonomicheskoj bezopasnosti Respubliki Belarus': teorija i praktika: materialy Respublikanskoj nauchno-prakticheskoj konferencii, Minsk, 19 dekabrja 2019 g. [Ensuring the Economic Security of the Republic of Belarus: Theory and Practice: Proceedings of the Republican Scientific-practical Conference, Minsk, 19 December 2019]. Minsk, Akademija Ministerstva vnutrennih del Respubliki Belarus', 2019, рр. 103–105 (In Russ.).</mixed-citation></citation-alternatives></ref><ref id="cit7"><label>7</label><citation-alternatives><mixed-citation xml:lang="ru">Fully convolutional networks for semantic segmentation / J. Long, E. Shelhamer, T. Darrell // Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), Boston, Massachusetts, USA, 7–12 June 2015. – Boston, 2015. – P. 3431–3440. – https://doi.org/10.1109/CVPR.2015.7298965.</mixed-citation><mixed-citation xml:lang="en">Long J., Shelhamer E., Darrell T. Fully convolutional networks for semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, Massachusetts, USA, 7–12 June 2015. Boston, 2015, рр. 3431–3440. https://doi.org/10.1109/CVPR.2015.7298965.</mixed-citation></citation-alternatives></ref><ref id="cit8"><label>8</label><citation-alternatives><mixed-citation xml:lang="ru">Ronneberger, O. U-Net: Convolutional networks for biomedical image segmentation / O. Ronneberger, P. Fischer, T. Brox // Medical Image Computing and Computer-Assisted Intervention – MICCAI 2015 : 18th Intern. Conf., Munich, 5–9 Oct. 2015 / ed.: N. Navab [et al.]. – Cham, 2015. – Pt. III. – P. 234–241. – (Lecture Notes in Computer Science ; vol. 9351). – https://doi.org/10.1007/978-3-319-24574-4_28.</mixed-citation><mixed-citation xml:lang="en">Ronneberger O., Fischer P., Brox T. U-Net: Convolutional networks for biomedical image segmentation. Medical Image Computing and Computer-Assisted Intervention – MICCAI 2015 : 18th International Conference, Munich, 5–9 October 2015. In N. Navab, J. Hornegger, W. M. Wells, A. F. Frangi (eds.). Cham, 2015, pt. III, рр. 234–241 (Lecture Notes in Computer Science; vol. 9351). https://doi.org/10.1007/978-3-319-24574-4_28.</mixed-citation></citation-alternatives></ref><ref id="cit9"><label>9</label><citation-alternatives><mixed-citation xml:lang="ru">Pyramid scene parsing network / H. Zhao, J. Shi, X. Qi [et al.] // Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), Honolulu, Hawaii, USA, 21–26 July 2017. – Honolulu, 2017. – P. 2881–2890. – https://doi.org/10.1109/CVPR.2017.660.</mixed-citation><mixed-citation xml:lang="en">H. Zhao, Shi J., Qi X., Wang X., Jia J. Pyramid scene parsing network. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, Hawaii, USA, 21–26 July 2017. Honolulu, 2017, рр. 2881–2890. https://doi.org/10.1109/CVPR.2017.660.</mixed-citation></citation-alternatives></ref><ref id="cit10"><label>10</label><citation-alternatives><mixed-citation xml:lang="ru">Encoder-decoder with atrous separable convolution for semantic image segmentation / L. C. Chen, Y. Zhu, G. Papandreou [et al.] // Proc. of the European Conf. on Computer Vision (ECCV), Munich, Germany, 8–14 Sept. 2018. – Munich, 2018. – P. 801–818. – https://doi.org/10.1007/978-3-030-01234-2_49.</mixed-citation><mixed-citation xml:lang="en">Chen L.-C., Papandreou G., Kokkinos I., Murphy K., Yuille A. L. Semantic image segmentation with deep convolutional nets and fully connected CRFs. International Conference on Learning Representations (ICLR), San Diego, CA, USA, 7–9 May 2015. San Diego, 2015, рр. 1–14. https://doi.org/10.48550/arXiv.1412.7062.</mixed-citation></citation-alternatives></ref><ref id="cit11"><label>11</label><citation-alternatives><mixed-citation xml:lang="ru">DeepLab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected CRFs / L. C. Chen, G. Papandreou, I. Kokkinos [et al.] // IEEE Transactions on Pattern Analysis and Machine Intelligence. – 2018. – Vol. 40, iss. 4. – P. 834–848. – https://doi.org/10.1109/TPAMI.2017.2699184.</mixed-citation><mixed-citation xml:lang="en">Chen L.-C., Papandreou G., Kokkinos I., Murphy K., Yuille A. L. DeepLab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected CRFs. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2018, vol. 40, iss. 4, рр. 834–848. https://doi.org/10.1109/TPAMI.2017.2699184.</mixed-citation></citation-alternatives></ref><ref id="cit12"><label>12</label><citation-alternatives><mixed-citation xml:lang="ru">Rethinking atrous convolution for semantic image segmentation / L.-C. Chen, G. Papandreou, F. Schroff, H. Adam. – 2017. – URL: https://arxiv.org/pdf/1706.05587 (date of access: 16.03.2026). – https://doi.org/10.48550/arXiv.1706.05587.</mixed-citation><mixed-citation xml:lang="en">Chen L.-C., Papandreou G., Schroff F., Adam H. Rethinking atrous convolution for semantic image segmentation, 2017. Available at: https://arxiv.org/pdf/1706.05587 (accessed: 16.03.2026). https://doi.org/10.48550/arXiv.1706.05587.</mixed-citation></citation-alternatives></ref><ref id="cit13"><label>13</label><citation-alternatives><mixed-citation xml:lang="ru">Encoder-decoder with atrous separable convolution for semantic image segmentation / L. C. Chen, Y. Zhu, G. Papandreou [et al.] // Proc. of the European Conf. on Computer Vision (ECCV), Munich, Germany, 8–14 Sept. 2018. – Munich, 2018. – P. 801–818. – https://doi.org/10.1007/978-3-030-01234-2_49.</mixed-citation><mixed-citation xml:lang="en">Chen L.-C., Zhu Y., Papandreou G., Schroff F., Adam H. Encoder-decoder with atrous separable convolution for semantic image segmentation. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018. Munich, 2018, рр. 801–818. https://doi.org/10.1007/978-3-030-01234-2_49.</mixed-citation></citation-alternatives></ref><ref id="cit14"><label>14</label><citation-alternatives><mixed-citation xml:lang="ru">An image is worth 16x16 words: Transformers for image recognition at scale / A. Dosovitskiy, L. Beyer, A. Kolesnikov [et al.] // Intern. Conf. on Learning Representations (ICLR), Vienna, Austria, 4 May 2021. – Vienna, 2021. – P. 1–21. – https://doi.org/10.48550/arXiv.2010.11929.</mixed-citation><mixed-citation xml:lang="en">Dosovitskiy A., Beyer L., Kolesnikov A., Weissenborn D., Zhai X., …, Houlsby N. An image is worth 16x16 words: Transformers for image recognition at scale. International Conference on Learning Representations (ICLR), Vienna, Austria, 4 May 2021. Vienna, 2021, рр. 1–21. https://doi.org/10.48550/arXiv.2010.11929.</mixed-citation></citation-alternatives></ref><ref id="cit15"><label>15</label><citation-alternatives><mixed-citation xml:lang="ru">Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers / S. Zheng, J. Lu, H. Zhao [et al.] // Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 20–25 June 2021. – Nashville, 2021. – P. 6881–6890. – https://doi.org/10.1109/CVPR46437.2021.00681.</mixed-citation><mixed-citation xml:lang="en">Zheng S., Lu J., Zhao H., Zhu X., Luo Z., Wang Y. Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 20–25 June 2021. Nashville, 2021, рр. 6881–6890. https://doi.org/10.1109/CVPR46437.2021.00681.</mixed-citation></citation-alternatives></ref><ref id="cit16"><label>16</label><citation-alternatives><mixed-citation xml:lang="ru">Attention mechanisms in semantic segmentation of remote sensing images / V. V. Godase, S. Takale, R. Ghodake, A. Mulani // Journal of Advancement in Electronics Signal Processing. – 2025. – Vol. 2. – P. 45–58.</mixed-citation><mixed-citation xml:lang="en">Godase V. V., Takale S., Ghodake R., Mulani A. Attention mechanisms in semantic segmentation of remote sensing images. Journal of Advancement in Electronics Signal Processing, 2025, vol. 2, рр. 45–58.</mixed-citation></citation-alternatives></ref><ref id="cit17"><label>17</label><citation-alternatives><mixed-citation xml:lang="ru">Hu, J. Squeeze-and-excitation networks / J. Hu, L. Shen, G. Sun // Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–22 June 2018. – Salt Lake City, 2018. – P. 7132–7141. – https://doi.org/10.1109/CVPR.2018.00745.</mixed-citation><mixed-citation xml:lang="en">Hu J., Shen L., Sun G. Squeeze-and-excitation networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–22 June 2018. Salt Lake City, 2018, рр. 7132–7141. https://doi.org/10.1109/CVPR.2018.00745.</mixed-citation></citation-alternatives></ref><ref id="cit18"><label>18</label><citation-alternatives><mixed-citation xml:lang="ru">CBAM: Convolutional block attention module / S. Woo, J. Park, J. Y. Lee, I. S. Kweon // Proc. of the European Conf. on Computer Vision (ECCV), Munich, Germany, 8–14 Sept. 2018. – Munich, 2018. – P. 3–19. – https://doi.org/10.1007/978-3-030-01234-2_1.</mixed-citation><mixed-citation xml:lang="en">Woo S., Park J., Lee J. Y., Kweon I. S. CBAM: Convolutional block attention module. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018. Munich, 2018, рр. 3–19. https://doi.org/10.1007/978-3-030-01234-2_1.</mixed-citation></citation-alternatives></ref><ref id="cit19"><label>19</label><citation-alternatives><mixed-citation xml:lang="ru">Hou, Q. Coordinate attention for efficient mobile network design / Q. Hou, D. Zhou, J. Feng // Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 20–25 June 2021. – Nashville, 2021. – P. 13713–13722. – https://doi.org/10.1109/CVPR46437.2021.01350.</mixed-citation><mixed-citation xml:lang="en">Hou Q., Zhou D., Feng J. Coordinate attention for efficient mobile network design. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 20–25 June 2021. Nashville, 2021, рр. 13713–13722. https://doi.org/10.1109/CVPR46437.2021.01350.</mixed-citation></citation-alternatives></ref><ref id="cit20"><label>20</label><citation-alternatives><mixed-citation xml:lang="ru">Big self-supervised models are strong semi-supervised learners / T. Chen, S. Kornblith, K. Swersky [et al.] // Advances in Neural Information Processing Systems (NeurIPS). – 2020. – Vol. 33. – P. 20735–20747. – https://doi.org/10.48550/arXiv.2006.10029.</mixed-citation><mixed-citation xml:lang="en">Chen T., Kornblith S., Swersky K., Norouzi M., Hinton G. Big self-supervised models are strong semi-supervised learners. Advances in Neural Information Processing Systems (NeurIPS), 2020, vol. 33, рр. 20735–20747. https://doi.org/10.48550/arXiv.2006.10029.</mixed-citation></citation-alternatives></ref><ref id="cit21"><label>21</label><citation-alternatives><mixed-citation xml:lang="ru">A comprehensive survey on contrastive learning / H. Hu, X. Wang, Y. Zhang [et al.] // Neurocomputing. – 2024. – Vol. 610. – Art. 128645. – URL: https://www.sciencedirect.com/science/article/abs/pii/S0925231224014164?via%3Dihub (date of access: 16.03.2026). – https://doi.org/10.1016/j.neucom.2024.128645.</mixed-citation><mixed-citation xml:lang="en">Hu H., Wang X., Zhang Y., Chen Q., Guan Q. A comprehensive survey on contrastive learning. Neurocomputing, 2024, vol. 610, art. 128645. Available at: https://www.sciencedirect.com/science/article/abs/pii/S0925231224014164?via%3Dihub (accessed 16.03.2026). https://doi.org/10.1016/j.neucom.2024.128645.</mixed-citation></citation-alternatives></ref><ref id="cit22"><label>22</label><citation-alternatives><mixed-citation xml:lang="ru">Supervised contrastive learning / P. Khosla, P. Teterwak, C. Wang [et al.] // Advances in Neural Information Processing Systems (NeurIPS). – 2020. – Vol. 33. – P. 18661–18673. – https://doi.org/10.48550/arXiv.2004.11362.</mixed-citation><mixed-citation xml:lang="en">Khosla P., Teterwak P., Wang C., Sarna A., Tian Y., …, Krishnan D. Supervised contrastive learning. Advances in Neural Information Processing Systems (NeurIPS), 2020, vol. 33, рр. 18661–18673. https://doi.org/10.48550/arXiv.2004.11362.</mixed-citation></citation-alternatives></ref><ref id="cit23"><label>23</label><citation-alternatives><mixed-citation xml:lang="ru">Deep residual learning for image recognition / K. He, X. Zhang, S. Ren, J. Sun // Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27–30 June 2016. – Las Vegas, 2016. – P. 770–778. – https://doi.org/10.1109/CVPR.2016.90.</mixed-citation><mixed-citation xml:lang="en">He K., Zhang X., Ren S., Sun J. Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27–30 June 2016. Las Vegas, 2016, рр. 770–778. https://doi.org/10.1109/CVPR.2016.90.</mixed-citation></citation-alternatives></ref><ref id="cit24"><label>24</label><citation-alternatives><mixed-citation xml:lang="ru">Comparative analysis of deep learning architectures for medical image classification / P.-N. Bui, D.-T. Le, J. Bum [et al.] // Bioengineering. – 2023. – Vol. 10, iss. 10. – Art. 1249. – URL: https://pmc.ncbi.nlm.nih.gov/articles/PMC10669434/ (date of access: 16.03.2026). – https://doi.org/10.3390/bioengineering10101249.</mixed-citation><mixed-citation xml:lang="en">Bui P.-N., Le D.-T., Bum J., Kim S., Song S. J., Choo H. Comparative analysis of deep learning architectures for medical image classification. Bioengineering, 2023, vol. 10, iss. 10, art. 1249. Available at: https://pmc.ncbi.nlm.nih.gov/articles/PMC10669434/ (accessed 16.03.2026). https://doi.org/10.3390/bioengineering10101249.</mixed-citation></citation-alternatives></ref><ref id="cit25"><label>25</label><citation-alternatives><mixed-citation xml:lang="ru">Kansal, K. ResNet-based deep learning approaches for histopathological breast cancer classification / K. Kansal, A. Singh, D. Srivastava, K. Kansal // 2025 Intern. Conf. on Artificial Intelligence and Emerging Technologies (ICAIET), Bhubaneswar, Odisha, India, 28–30 Aug. 2025. – Bhubaneswar, 2025. – P. 1–4. – https://doi.org/10.1109/ICAIET.2025.11210953.</mixed-citation><mixed-citation xml:lang="en">Kansal K., Singh A., Srivastava D., Kansal K. ResNet-based deep learning approaches for histopathological breast cancer classification. 2025 International Conference on Artificial Intelligence and Emerging Technologies (ICAIET), Bhubaneswar, Odisha, India, 28–30 August 2025. Bhubaneswar, 2025, рр. 1–4. https://doi.org/10.1109/ICAIET.2025.11210953.</mixed-citation></citation-alternatives></ref><ref id="cit26"><label>26</label><citation-alternatives><mixed-citation xml:lang="ru">Kordemir, M. A mask R-CNN approach for detection and classification of brain tumours from MR images / M. Kordemir, K. K. Cevik, A. Bozkurt // Computer Methods in Biomechanics and Biomedical Engineering: Imaging &amp; Visualization. – 2024. – Vol. 11, iss. 7. – URL: https://www.tandfonline.com/doi/full/10.1080/21681163.2023.2301391 (date of access: 16.03.2026). – https://doi.org/10.1080/21681163.2023.2301391.</mixed-citation><mixed-citation xml:lang="en">Kordemir M., Cevik K. K., Bozkurt A. A mask R-CNN approach for detection and classification of brain tumours from MR images. Computer Methods in Biomechanics and Biomedical Engineering: Imaging &amp; Visualization, 2024, vol. 11, iss. 7. Available at: https://www.tandfonline.com/doi/full/10.1080/21681163.2023.2301391 (accessed 16.03.2026). https://doi.org/10.1080/21681163.2023.2301391.</mixed-citation></citation-alternatives></ref><ref id="cit27"><label>27</label><citation-alternatives><mixed-citation xml:lang="ru">A comparative analysis of attention mechanisms in 3D CNN-based hyperspectral image super-resolution / M. S. Zitouni, M. Q. Alkhatib, N. Aburaed, H. Ahmad // 2024 14th Workshop on Hyperspectral Imaging and Signal Processing: Evolution in Remote Sensing (WHISPERS), Helsinki, Finland, 09–11 Dec. 2024. – Helsinki, 2024. – P. 1–5. – https://doi.org/10.1109/WHISPERS.2024.10876507.</mixed-citation><mixed-citation xml:lang="en">Zitouni M. S., Alkhatib M. Q., Aburaed N., Ahmad H. A comparative analysis of attention mechanisms in 3D CNN-based hyperspectral image super-resolution. 2024 14th Workshop on Hyperspectral Imaging and Signal Processing: Evolution in Remote Sensing (WHISPERS), Helsinki, Finland, 09–11 December 2024. Helsinki, 2024, рр. 1–5. https://doi.org/10.1109/WHISPERS.2024.10876507.</mixed-citation></citation-alternatives></ref><ref id="cit28"><label>28</label><citation-alternatives><mixed-citation xml:lang="ru">A new semantic segmentation method for remote sensing images integrating coordinate attention and SPD-Conv / Z. Yang, Q. Wu, F. Zhang [et al.] // Symmetry. – 2023. – Vol. 15, iss. 5. – Art. 1037. – URL: https://www.mdpi.com/2073-8994/15/5/1037 (date of access: 16.03.2026). – https://doi.org/10.3390/sym15051037.</mixed-citation><mixed-citation xml:lang="en">Yang Z., Wu Q., Zhang F., Zhang X., Chen X., Gao Y. A new semantic segmentation method for remote sensing images integrating coordinate attention and SPD-Conv. Symmetry, 2023, vol. 15, iss. 5, art. 1037. Available at: https://www.mdpi.com/2073-8994/15/5/1037 (accessed 16.03.2026). https://doi.org/10.3390/sym15051037.</mixed-citation></citation-alternatives></ref><ref id="cit29"><label>29</label><citation-alternatives><mixed-citation xml:lang="ru">Li, Z. A comparative study of Contrastive Self-Supervised Learning (CSSL): Methods, technologies, and applications / Z. Li // Transactions on Computer Science and Intelligent Systems Research. – 2025. – Vol. 10. – P. 92–97. – https://doi.org/10.62051/2hgr2j26.</mixed-citation><mixed-citation xml:lang="en">Li Z. A comparative study of Contrastive Self-Supervised Learning (CSSL): Methods, technologies, and applications. Transactions on Computer Science and Intelligent Systems Research, 2025, vol. 10, рр. 92–97. https://doi.org/10.62051/2hgr2j26.</mixed-citation></citation-alternatives></ref><ref id="cit30"><label>30</label><citation-alternatives><mixed-citation xml:lang="ru">Wang, Y. Multi-label guided supervised contrastive learning for earth observation pretraining / Y. Wang, C. M. Albrecht, X. X. Zhu // 2024 IEEE Intern. Geoscience and Remote Sensing Symp. (IGARSS), Athens, Greece, 7–12 July 2024. – Athens, 2024. – P. 7568–7571. – https://doi.org/10.1109/IGARSS53475.2024.10642289.</mixed-citation><mixed-citation xml:lang="en">Wang Y., Albrecht C. M., Zhu X. X. Multi-label guided supervised contrastive learning for earth observation pretraining. 2024 IEEE International Geoscience and Remote Sensing Symposium (IGARSS), Athens, Greece, 7–12 July 2024. Athens, 2024, рр. 7568–7571. https://doi.org/10.1109/IGARSS53475.2024.10642289.</mixed-citation></citation-alternatives></ref><ref id="cit31"><label>31</label><citation-alternatives><mixed-citation xml:lang="ru">Searching for MobileNetV3 / A. Howard, M. Sandler, B. Chen [et al.] // Proc. of the IEEE/CVF Intern. Conf. on Computer Vision (ICCV), Seoul, Korea (South), 27 Oct. – 02 Nov. 2019. – Seoul, 2019. – P. 1314–1324. – https://doi.org/10.1109/ICCV.2019.00140.</mixed-citation><mixed-citation xml:lang="en">Howard A., Sandler M., Chen B., Wang W., Chen L.-C., Tan M. Searching for MobileNetV3. Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Korea (South), 27 October – 02 November 2019. Seoul, 2019, рр. 1314–1324. https://doi.org/10.1109/ICCV.2019.00140.</mixed-citation></citation-alternatives></ref><ref id="cit32"><label>32</label><citation-alternatives><mixed-citation xml:lang="ru">FiLM: Visual reasoning with a general conditioning layer / E. Perez, F. Strub, H. de Vries [et al.] // Proc. of the AAAI Conf. on Artificial Intelligence, New Orleans, Louisiana, USA, 2–7 Febr. 2018. – New Orleans, 2018. – Vol. 32, iss. 1. – P. 3942–3951. – https://doi.org/10.1609/aaai.v32i1.11823.</mixed-citation><mixed-citation xml:lang="en">Perez E., Strub F., Vries H. de, Dumoulin V., Courville A. FiLM: Visual reasoning with a general conditioning layer. Proceedings of the AAAI Conference on Artificial Intelligence, New Orleans, Louisiana, USA, 2–7 February 2018. New Orleans, 2018, vol. 32, iss. 1, рр. 3942–3951. https://doi.org/10.1609/aaai.v32i1.11823.</mixed-citation></citation-alternatives></ref></ref-list><fn-group><fn fn-type="conflict"><p>The authors declare that there are no conflicts of interest present.</p></fn></fn-group></back></article>
