A mathematical model for the automatic extraction of idioms in Turkic languages

Authors

  • E.Sh. Nazirova Tashkent University of Information Technologies named after Muhammad al-Khwarizmi Author
  • N.B. Madetbaeva Digital Technologies and Artificial Intelligence Development Research Institute Author
  • F.Sh. Mirzaeva Tashkent University of Information Technologies named after Muhammad al-Khwarizmi Author
  • M.B. Islomova Inha University in Tashkent Author

DOI:

https://doi.org/10.71310/pcam.4_74.2026.10

Keywords:

phraseological unit, idiom detection, binary classification, pointwise mutual information, syntactic fixedness, Uzbek language

Abstract

This paper presents a complete mathematical model for the automatic extraction of phraseological units (idioms) from texts in Turkic languages, with a focus on Uzbek. The task is formulated as a binary classification problem and is solved by a three-stage pipeline. At the first stage, a set of idiom candidates is generated by a linguistic filter based on verb-headed grammatical templates. At the second stage, each candidate is mapped to a feature vector that combines statistical association (pointwise mutual information), syntactic fixedness (a flexibility index), semantic non-compositionality, and template-lexical indicators. At the third stage, a decision function is trained using four classification algorithms, namely Naive Bayes, logistic regression, support vector machines (SVM), and gradient boosting. The model was evaluated on a combined Uzbek corpus of 1.07 million lemma-tokens, supported by a phraseological ontology of 1,362 units and a manually annotated, stratified gold-standard set of 400 sentences. The best overall result was achieved by the SVM (????1 = 0.795; recall = 0.906), followed closely by gradient boosting (????1 = 0.791), whereas the lexical string-matching baseline proved nearly ineffective (????1 = 0.100) and ontology-based matching provided the highest precision (???? = 0.854). Stratified analysis showed that the main advantage of the model lies in recognizing morphologically transformed forms of idioms, where the learning-based model reached a recall of 0.919 against 0.041 for the lexical baseline. The role of each method within the processing pipeline and its most effective area of application are discussed.

References

Sag I. A., Baldwin T., Bond F., Copestake A., Flickinger D. 2002. Multiword expressions: A pain in the neck for NLP. Computational Linguistics and Intelligent Text Processing (CICLing 2002). LNCS, Vol. 2276. Berlin: Springer. 1–15.

Baldwin T., Kim S. N. 2010. Multiword expressions. Handbook of Natural Language Processing. 2nd ed. Boca Raton: CRC Press. 267–292.

Ramisch C. 2015. Multiword expressions acquisition: A generic and open framework. Cham: Springer. 230 p.

Constant M., Eryiğit G., Monti J., van der Plas L., Ramisch C., Rosner M., Todirascu A. 2017. Multiword expression processing: A survey. Computational Linguistics. 43(4): 837–892.

Church K. W., Hanks P. 1990. Word association norms, mutual information, and lexicography. Computational Linguistics. 16(1): 22–29.

Manning C. D., Schütze H. 1999. Foundations of statistical natural language processing. Cambridge: MIT Press. 680 p.

Evert S. 2005. The statistics of word cooccurrences: Word pairs and collocations. PhD dissertation. Stuttgart: Universität Stuttgart. 353 p.

Pecina P. 2010. Lexical association measures and collocation extraction. Language Resources and Evaluation. 44(1–2): 137–158.

Frantzi K., Ananiadou S., Mima H. 2000. Automatic recognition of multi-word terms: the C-value/NC-value method. International Journal on Digital Libraries. 3(2): 115–130.

Fazly A., Cook P., Stevenson S. 2009. Unsupervised type and token identification of idiomatic expressions. Computational Linguistics. 35(1): 61–103.

Cook P., Fazly A., Stevenson S. 2008. The VNC-Tokens dataset. Proceedings of the LREC Workshop Towards a Shared Task for Multiword Expressions (MWE 2008). Marrakech. 19–22.

Katz G., Giesbrecht E. 2006. Automatic identification of non-compositional multi-word expressions using latent semantic analysis. Proceedings of the Workshop on Multiword Expressions. Sydney: ACL. 12–19.

Salehi B., Cook P., Baldwin T. 2015. A word embedding approach to predicting the compositionality of multiword expressions. Proceedings of NAACL-HLT 2015. Denver: ACL. 977–983.

Cordeiro S., Villavicencio A., Idiart M., Ramisch C. 2019. Unsupervised compositionality prediction of nominal compounds. Computational Linguistics. 45(1): 1–57.

Mikolov T., Chen K., Corrado G., Dean J. 2013. Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781.

Devlin J., Chang M.-W., Lee K., Toutanova K. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. Proceedings of NAACL-HLT 2019. Minneapolis: ACL. 4171–4186.

Reimers N., Gurevych I. 2019. Sentence-BERT: Sentence embeddings using Siamese BERTnetworks. Proceedings of EMNLP-IJCNLP 2019. Hong Kong: ACL. 3982–3992.

Feng F., Yang Y., Cer D., Arivazhagan N., Wang W. 2022. Language-agnostic BERT sentence embedding. Proceedings of the 60th Annual Meeting of the ACL. Dublin: ACL. 878–891.

Zeng Z., Bhat S. 2021. Idiomatic expression identification using semantic compatibility. Transactions of the Association for Computational Linguistics. 9: 1546–1562.

Tayyar Madabushi H., Gow-Smith E., Garcia M., Scarton C., Idiart M., Villavicencio A. 2022. SemEval-2022 Task 2: Multilingual idiomaticity detection and sentence embedding. Proceedings of SemEval-2022. Seattle: ACL. 107–121.

Savary A., Ramisch C., Cordeiro S., Sangati F., Vincze V., QasemiZadeh B., Candito M., Cap F., Giouli V., Stoyanova I., Doucet A. 2017. The PARSEME shared task on automatic identification of verbal multiword expressions. Proceedings of the 13th Workshop on Multiword Expressions (MWE 2017). Valencia: ACL. 31–47.

Ramisch C., Cordeiro S., Savary A. [et al.] 2018. Edition 1.1 of the PARSEME shared task on automatic identification of verbal multiword expressions. Proceedings of LAW-MWECxG-2018. Santa Fe: ACL. 222–240.

Berk G., Erden B., Güngör T. 2018. Deep-BGT at PARSEME shared task 2018: Bidirectional LSTM-CRF model for verbal multiword expression identification. Proceedings of LAW-MWE-CxG-2018. Santa Fe: ACL. 248–253.

Oflazer K. 1994. Two-level description of Turkish morphology. Literary and Linguistic Computing. 9(2): 137–148.

Oflazer K., Saraçlar M. (eds.) 2018. Turkish natural language processing. Cham: Springer. 367 p.

Kuriyozov E., Doval Y., Gómez-Rodríguez C. 2020. Cross-lingual word embeddings for Turkic languages. Proceedings of LREC 2020. Marseille: ELRA. 4054–4062.

Salaev U. 2024. UzMorphAnalyser: A morphological analysis model for the Uzbek language using inflectional endings. AIP Conference Proceedings. 3244: Art. 030058. doi: https://doi.org/10.1063/5.0241461.

Salaev U., Kuriyozov E., Gómez-Rodríguez C. 2022. SimRelUz: Similarity and relatedness scores as a semantic evaluation dataset for Uzbek language. Proceedings of SIGUL 2022. Marseille: ELRA. 199–206.

Vinogradov V. V. 1977. Ob osnovnykh tipakh frazeologicheskikh edinits v russkom yazyke [On the main types of phraseological units in the Russian language]. Vinogradov V. V. Izbrannye trudy. Leksikologiya i leksikografiya [Selected works. Lexicology and lexicography]. Moscow: Nauka. 140–161. (In Russian)

Shanskiy N. M. 1985. Frazeologiya sovremennogo russkogo yazyka [Phraseology of the modern Russian language]. 3rd ed. Moscow: Vysshaya shkola. 160 p. (In Russian)

Rahmatullayev Sh. 1966. O‘zbek frazeologiyasining ba’zi masalalari [Some issues of Uzbek phraseology]. Tashkent: Fan. (In Uzbek)

Rahmatullayev Sh. 1978. O‘zbek tilining izohli frazeologik lug‘ati [Explanatory phraseological dictionary of the Uzbek language]. Tashkent: O‘qituvchi. (In Uzbek)

Cortes C., Vapnik V. 1995. Support-vector networks. Machine Learning. 20(3): 273–297.

Friedman J. H. 2001. Greedy function approximation: A gradient boosting machine. The Annals of Statistics. 29(5): 1189–1232.

Chen T., Guestrin C. 2016. XGBoost: A scalable tree boosting system. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. San Francisco: ACM. 785–794.

Hastie T., Tibshirani R., Friedman J. 2009. The elements of statistical learning. 2nd ed. New York: Springer. 745 p.

Fawcett T. 2006. An introduction to ROC analysis. Pattern Recognition Letters. 27(8): 861–874.

Pedregosa F., Varoquaux G., Gramfort A. [et al.] 2011. Scikit-learn: Machine learning in Python. Journal of Machine Learning Research. 12: 2825–2830.

Downloads

Published

2026-09-15

Issue

Section

Статьи