A three-vector algorithm for detecting Russian-language neural-network-generated text fragments based on probabilistic analysis
https://doi.org/10.37661/1/1816-0301-2026-23-3-76-91
Abstract
Objectives. The purpose of the work is to develop and programmatically implement a three-vector algorithm for detecting Russian-language neural network text fragments based on probabilistic analysis. The object of the study is Russian-language texts, the subject is statistical signs of their origin. Special attention is paid to the quantitative comparison of the proposed approach with the basic single-vector detection methods the perplexy detector and the GLTR rank analysis method.
Methods. The proposed algorithm combines three feature vectors: token predictability, the proportion of rare lexemes, and the information density of the text, estimated via the algorithmic compression ratio. The author's regression language model rugpt3small is used as a probabilistic core. The implementation is made in Python using the transformers, PyTorch, NLTK, and zlib libraries. The experimental validation was conducted on a sample of 100 mixed documents (16,000 sentences) of five thematic groups, the neural network fragments of which were generated by four language models DeepSeek, GPT, Qwen and YandexAI. The quality was assessed by the deviation from the reference proportion of the neural network text and by classification metrics.
Results. After applying the calibration adjustment, the average index for the sample was 45.2 points against a reference level of 50, and the values for all four models were in the narrow range of 44.445.9 points, which indicates the algorithm's stability to the choice of a generative model. The deviation of the average index from the standard was 4.8 points versus 11.1 points for the method based on perplexity and 7.6 points for GLTR; the F1 measure at the sentence level reached 0.76 versus 0.63 and 0.66, respectively.
Conclusion. The three-vector algorithm has proven its effectiveness in the task of detecting neural network fragments and can serve as the basis for tools for verifying academic texts and anti-plagiarism systems. Further development of the research is related to the expansion of the sample, testing the resistance to paraphrasing and editing, as well as conducting ablative analysis
About the Authors
K. S. KrezBelarus
Karyna S. Krez , Postgraduate Student, Assistant of the Department of Information and Computer Systems Design
st. P. Brovki, 6, Minsk, 220013
Y. N. Shneiderov
Belarus
Yevgeny N. Shneiderov , Cand. Sci. (Eng.), Assoc. Prof., Assoc. Prof. of the Department of Information and Computer Systems Design, Vice-Rector for Academic Affairs
st. P. Brovki, 6, Minsk, 220013
I. P. Galyak
Belarus
Ivan P. Galyak , Student at the Department of Information and Computer Systems Design
st. P. Brovki, 6, Minsk, 220013
D. D. Efimenko
Daniil D. Efimenko , Student at the Department of Information and Computer Systems Design
st. P. Brovki, 6, Minsk, 220013
References
1. Ivakhnenko E. N., Nikolsky V. S. ChatGPT in higher education and science: A threat or a valuable resource? Vysshee obrazovanie v Rossii [Higher Education in Russia], 2023, vol. 32, no. 4, рр. 9–22 (In Russ.).
2. Fedotova A. M., Romanov A. S. Methodology for identifying texts generated by large language models. Informatika i avtomatizatsiya [Computer Science and Automation], 2025, vol. 24, no. 5, рр. 1444–1470 (In Russ.).
3. Machkovskaya L. Ya., Fatina A. V., Vetoshkina O. A. Generative artificial intelligence models in teaching Russian as a foreign language: Opportunities, limitations, and risks. Mezhdunarodnyy zhurnal gumanitarnykh i estestvennykh nauk [International Journal of Humanities and Natural Sciences], 2025, no. 11-1, pp. 73–78 (In Russ.).
4. Aydagulova A. R. Features of texts generated by artificial intelligence. Vestnik Bashkirskogo gosudarstvennogo pedagogicheskogo universiteta im. M. Akmully [Bulletin of M. Akmulla Bashkir State Pedagogical University], 2023, no. 4 (72), рр. 154–156 (In Russ.).
5. Bryzgalina E. V. Challenges of artificial intelligence technologies for the ethical review of living systems research. Teoreticheskaya i prikladnaya etika: traditsii i perspektivy: materialy XVI Mezhdunarodnoj konferencii, Sankt-Peterburg, 16–18 nojabrja 2023 g. [Theoretical and Applied Ethics: Traditions and Prospects: Proceedings of the 16th International Conference, Saint Petersburg, 16–18 November 2023]. Saint Petersburg, Izdatel'stvo Sankt-Peterburgskogo gosudarstvennogo universiteta, 2024, pp. 57–58 (In Russ.).
6. Veys I. A. Recognition of artificial intelligence application patterns in text creation. Sovremennye innovatsii, sistemy i tekhnologii [Modern Innovations, Systems and Technologies], 2025, vol. 5, no. 1, рр. 1033–1040 (In Russ.).
7. Vasilevskaya A. V. AI in quantitative research: Testing resilience to manipulation and factual distortions. Teoreticheskaya i prikladnaya etika: traditsii i perspektivy: materialy XVI Mezhdunarodnoj konferencii, Sankt-Peterburg, 16–18 nojabrja 2023 g. [Theoretical and Applied Ethics: Traditions and Prospects: Proceedings of the 16th International Conference, Saint Petersburg, 16–18 November 2023]. Saint Petersburg, Izdatel'stvo Sankt-Peterburgskogo gosudarstvennogo universiteta, 2024, рр. 67–68 (In Russ.).
8. Nikitina A. S. Human or AI: On the issue of authorship. Vestnik molodykh uchenykh i spetsialistov Samarskogo universiteta [Bulletin of Young Scientists and Specialists of Samara University], 2024, no. 1 (24), рр. 182–186 (In Russ.).
9. Gehrmann S., Strobelt H., Rush A. M. GLTR: Statistical detection and visualization of generated text. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, Florence, Italy, 28 July – 2 August 2019. Florence, 2019, рр. 111–116.
10. Mitchell E., Lee Y., Khazatsky A., Manning C. D., Finn C. DetectGPT: Zero-shot machine-generated text detection using probability curvature. Proceedings of the 40th International Conference on Machine Learning (ICML 2023), Honolulu, Hawaii, USA, 23–29 July 2023. Honolulu, 2023, рр. 24950–24962.
11. Kirchenbauer J., Geiping J., Wen Y., Katz J., Miers I., Goldstein T. A watermark for large language models. Proceedings of the 40th International Conference on Machine Learning (ICML 2023), Honolulu, Hawaii, USA, 23–29 July 2023. Honolulu, 2023, рр. 17061–17084.
Review
For citations:
Krez K.S., Shneiderov Y.N., Galyak I.P., Efimenko D.D. A three-vector algorithm for detecting Russian-language neural-network-generated text fragments based on probabilistic analysis. Informatics. 2026;23(3):76-91. (In Russ.) https://doi.org/10.37661/1/1816-0301-2026-23-3-76-91
JATS XML

















