Математическая модель автоматического выделения фразеологизмов в тюркских языках
DOI:
https://doi.org/10.71310/pcam.4_74.2026.10Ключевые слова:
фразеологизм, автоматическое выделение фразеологизмов, бинарная классификация, поточечная взаимная информация, синтаксическая устойчивость, узбекский языкАннотация
В статье представлена полная математическая модель автоматического выделения фразеологизмов из текстов на тюркских языках, прежде всего на узбекском. Задача поставлена как задача бинарной классификации и решается трёхэтапной цепочкой. На первом этапе лингвистический фильтр, опирающийся на глагольные грамматические шаблоны, формирует множество кандидатов. На втором этапе каждый кандидат представляется вектором признаков, объединяющим статистическую связанность компонентов (поточечную взаимную информацию), синтаксическую устойчивость (индекс гибкости), семантическую некомпозициональность и шаблонно-лексические индикаторы. На третьем этапе решающая функция обучается четырьмя алгоритмами классификации: наивным байесовским классификатором, логистической регрессией, методом опорных векторов (SVM) и градиентным бустингом. Модель испытана на объединённом узбекском корпусе объёмом 1,07 млн лемма-токенов с опорой на фразеологическую онтологию из 1362 единиц и вручную размеченный стратифицированный эталонный набор из 400 предложений. Наилучший общий результат показал метод опорных векторов (????1 = 0,795; полнота 0,906), градиентный бустинг уступил ему минимально (????1 = 0,791), лексическое сопосталение оказалось практически неэффективным (????1 = 0,100), а онтологическое сопоставление обеспечило наивысшую точность (???? = 0,854). Стратифицированный анализ показал, что основное преимущество модели проявляется в распознавании морфологически изменённых форм фразеологизмов. Обсуждается место каждого метода в цепочке обработки и область его эффективного применения.
Библиографические ссылки
Sag I. A., Baldwin T., Bond F., Copestake A., Flickinger D. 2002. Multiword expressions: A pain in the neck for NLP. Computational Linguistics and Intelligent Text Processing (CICLing 2002). LNCS, Vol. 2276. Berlin: Springer. 1–15.
Baldwin T., Kim S. N. 2010. Multiword expressions. Handbook of Natural Language Processing. 2nd ed. Boca Raton: CRC Press. 267–292.
Ramisch C. 2015. Multiword expressions acquisition: A generic and open framework. Cham: Springer. 230 p.
Constant M., Eryiğit G., Monti J., van der Plas L., Ramisch C., Rosner M., Todirascu A. 2017. Multiword expression processing: A survey. Computational Linguistics. 43(4): 837–892.
Church K. W., Hanks P. 1990. Word association norms, mutual information, and lexicography. Computational Linguistics. 16(1): 22–29.
Manning C. D., Schütze H. 1999. Foundations of statistical natural language processing. Cambridge: MIT Press. 680 p.
Evert S. 2005. The statistics of word cooccurrences: Word pairs and collocations. PhD dissertation. Stuttgart: Universität Stuttgart. 353 p.
Pecina P. 2010. Lexical association measures and collocation extraction. Language Resources and Evaluation. 44(1–2): 137–158.
Frantzi K., Ananiadou S., Mima H. 2000. Automatic recognition of multi-word terms: the C-value/NC-value method. International Journal on Digital Libraries. 3(2): 115–130.
Fazly A., Cook P., Stevenson S. 2009. Unsupervised type and token identification of idiomatic expressions. Computational Linguistics. 35(1): 61–103.
Cook P., Fazly A., Stevenson S. 2008. The VNC-Tokens dataset. Proceedings of the LREC Workshop Towards a Shared Task for Multiword Expressions (MWE 2008). Marrakech. 19–22.
Katz G., Giesbrecht E. 2006. Automatic identification of non-compositional multi-word expressions using latent semantic analysis. Proceedings of the Workshop on Multiword Expressions. Sydney: ACL. 12–19.
Salehi B., Cook P., Baldwin T. 2015. A word embedding approach to predicting the compositionality of multiword expressions. Proceedings of NAACL-HLT 2015. Denver: ACL. 977–983.
Cordeiro S., Villavicencio A., Idiart M., Ramisch C. 2019. Unsupervised compositionality prediction of nominal compounds. Computational Linguistics. 45(1): 1–57.
Mikolov T., Chen K., Corrado G., Dean J. 2013. Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781.
Devlin J., Chang M.-W., Lee K., Toutanova K. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. Proceedings of NAACL-HLT 2019. Minneapolis: ACL. 4171–4186.
Reimers N., Gurevych I. 2019. Sentence-BERT: Sentence embeddings using Siamese BERTnetworks. Proceedings of EMNLP-IJCNLP 2019. Hong Kong: ACL. 3982–3992.
Feng F., Yang Y., Cer D., Arivazhagan N., Wang W. 2022. Language-agnostic BERT sentence embedding. Proceedings of the 60th Annual Meeting of the ACL. Dublin: ACL. 878–891.
Zeng Z., Bhat S. 2021. Idiomatic expression identification using semantic compatibility. Transactions of the Association for Computational Linguistics. 9: 1546–1562.
Tayyar Madabushi H., Gow-Smith E., Garcia M., Scarton C., Idiart M., Villavicencio A. 2022. SemEval-2022 Task 2: Multilingual idiomaticity detection and sentence embedding. Proceedings of SemEval-2022. Seattle: ACL. 107–121.
Savary A., Ramisch C., Cordeiro S., Sangati F., Vincze V., QasemiZadeh B., Candito M., Cap F., Giouli V., Stoyanova I., Doucet A. 2017. The PARSEME shared task on automatic identification of verbal multiword expressions. Proceedings of the 13th Workshop on Multiword Expressions (MWE 2017). Valencia: ACL. 31–47.
Ramisch C., Cordeiro S., Savary A. [et al.] 2018. Edition 1.1 of the PARSEME shared task on automatic identification of verbal multiword expressions. Proceedings of LAW-MWECxG-2018. Santa Fe: ACL. 222–240.
Berk G., Erden B., Güngör T. 2018. Deep-BGT at PARSEME shared task 2018: Bidirectional LSTM-CRF model for verbal multiword expression identification. Proceedings of LAW-MWE-CxG-2018. Santa Fe: ACL. 248–253.
Oflazer K. 1994. Two-level description of Turkish morphology. Literary and Linguistic Computing. 9(2): 137–148.
Oflazer K., Saraçlar M. (eds.) 2018. Turkish natural language processing. Cham: Springer. 367 p.
Kuriyozov E., Doval Y., Gómez-Rodríguez C. 2020. Cross-lingual word embeddings for Turkic languages. Proceedings of LREC 2020. Marseille: ELRA. 4054–4062.
Salaev U. 2024. UzMorphAnalyser: A morphological analysis model for the Uzbek language using inflectional endings. AIP Conference Proceedings. 3244: Art. 030058. doi: https://doi.org/10.1063/5.0241461.
Salaev U., Kuriyozov E., Gómez-Rodríguez C. 2022. SimRelUz: Similarity and relatedness scores as a semantic evaluation dataset for Uzbek language. Proceedings of SIGUL 2022. Marseille: ELRA. 199–206.
Vinogradov V. V. 1977. Ob osnovnykh tipakh frazeologicheskikh edinits v russkom yazyke [On the main types of phraseological units in the Russian language]. Vinogradov V. V. Izbrannye trudy. Leksikologiya i leksikografiya [Selected works. Lexicology and lexicography]. Moscow: Nauka. 140–161. (In Russian)
Shanskiy N. M. 1985. Frazeologiya sovremennogo russkogo yazyka [Phraseology of the modern Russian language]. 3rd ed. Moscow: Vysshaya shkola. 160 p. (In Russian)
Rahmatullayev Sh. 1966. O‘zbek frazeologiyasining ba’zi masalalari [Some issues of Uzbek phraseology]. Tashkent: Fan. (In Uzbek)
Rahmatullayev Sh. 1978. O‘zbek tilining izohli frazeologik lug‘ati [Explanatory phraseological dictionary of the Uzbek language]. Tashkent: O‘qituvchi. (In Uzbek)
Cortes C., Vapnik V. 1995. Support-vector networks. Machine Learning. 20(3): 273–297.
Friedman J. H. 2001. Greedy function approximation: A gradient boosting machine. The Annals of Statistics. 29(5): 1189–1232.
Chen T., Guestrin C. 2016. XGBoost: A scalable tree boosting system. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. San Francisco: ACM. 785–794.
Hastie T., Tibshirani R., Friedman J. 2009. The elements of statistical learning. 2nd ed. New York: Springer. 745 p.
Fawcett T. 2006. An introduction to ROC analysis. Pattern Recognition Letters. 27(8): 861–874.
Pedregosa F., Varoquaux G., Gramfort A. [et al.] 2011. Scikit-learn: Machine learning in Python. Journal of Machine Learning Research. 12: 2825–2830.
Загрузки
Опубликован
Выпуск
Раздел
Лицензия
Copyright (c) 2026 Э.Ш. Назирова

Это произведение доступно по лицензии Creative Commons «Attribution» («Атрибуция») 4.0 Всемирная.