A semantic linking model for Turkic languages: a fairness-constrained cross-lingual approach for low-resource settings
DOI:
https://doi.org/10.71310/pcam.4_74.2026.11Keywords:
Turkic languages, cross-lingual embeddings, low-resource languages, noisy supervision, fairness constraint, Procrustes problemAbstract
We propose a semantic linking model for the Turkic language family, most members of which are considered low-resource from the computational linguistics viewpoint. The theoretical foundation of the model is Liddy’s layered scheme of natural language processing, compressed here into four levels: graphemic-morphological, syntactic, semantic, and cross-lingual. Rather than treating expert-built knowledge resources as hard constraints, we introduce them as a noisy supervision signal weighted by a per-language reliability coefficient ????????; this construction subsumes both classical retrofitting (???????? → 1) and purely distributional models (???????? → 0) as special cases of a single scheme. The optimization problem is augmented with a fairness constraint bounding the gap between the best and worst per-language quality, which prevents the model from drifting toward resource-rich languages at the expense of low-resource ones. A two-level algorithm is derived that combines Adam updates for embeddings, a closed-form orthogonal Procrustes step for cross-lingual mappings, and softmax reparametrisation for the language-weight simplex with a projected subgradient step for the dual variable; a standard ????(1/√????) bound is stated for the outer loop. The evaluation protocol includes five baselines that isolate the contribution of each model component, intrinsic metrics (Spearman correlation on SimRelUz, Precision@1 and Precision@5 for bilingual lexicon induction) and extrinsic metrics (chrF++ for downstream machine translation), with 95% bootstrap confidence intervals and Holm–Bonferroni correction for multiple comparisons.
References
Agostini A. [et al.] 2021. UZWORDNET: A lexical-semantic database for the Uzbek language. Proceedings of the 11th Global WordNet Conference (GWC-2021). Global WordNet Association. 8–19.
Artetxe M., Labaka G., Agirre E. 2018. A robust self-learning method for fully unsupervised cross-lingual mappings of word embeddings. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (ACL 2018). Vol. 1. 789–798.
Batsuren K., Bella G., Giunchiglia F. 2022. A large and evolving cognate database. Language Resources and Evaluation. 56: 165–189. doi: https://doi.org/10.1007/s10579-021-09544-6.
Bojanowski P., Grave E., Joulin A., Mikolov T. 2017. Enriching word vectors with subword information. Transactions of the Association for Computational Linguistics. 5: 135–146.
Church K., Yuan X., Guo S., Wu Z., Yang Y., Chen Z. 2021. Emerging trends: A gentle introduction to fine-tuning. Natural Language Engineering. 27(6): 763–778.
Collobert R., Weston J., Bottou L., Karlen M., Kavukcuoglu K., Kuksa P. 2011. Natural language processing (almost) from scratch. Journal of Machine Learning Research. 12: 2493–2537.
Conneau A., Khandelwal K., Goyal N., Chaudhary V., Wenzek G., Guzmán F., Grave E., Ott M., Zettlemoyer L., Stoyanov V. 2020. Unsupervised cross-lingual representation learning at scale. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL 2020). 8440–8451.
Jurafsky D., Martin J. H. 2025. Speech and language processing: An introduction to natural language processing, computational linguistics, and speech recognition with language models. 3rd ed. Online manuscript. Available at: https://web.stanford.edu/~jurafsky/slp3.
Liddy E. D. 2001. Natural language processing. Encyclopedia of Library and Information Science. 2nd ed. New York: Marcel Dekker.
Mann W. C., Thompson S. A. 1988. Rhetorical structure theory: Toward a functional theory of text organization. Text. 8(3): 243–281.
Manning C. D., Schütze H. 1999. Foundations of statistical natural language processing. Cambridge, MA: MIT Press. 680 p.
Mikolov T., Chen K., Corrado G., Dean J. 2013. Efficient estimation of word representations in vector space. Proceedings of the International Conference on Learning Representations (ICLR 2013), Workshop Track. arXiv:1301.3781.
Miller G. A. 1995. WordNet: A lexical database for English. Communications of the ACM. 38(11): 39–41.
Pennington J., Socher R., Manning C. D. 2014. GloVe: Global vectors for word representation. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP). 1532–1543.
Ruder S., Vulić I., Søgaard A. 2019. A survey of cross-lingual word embedding models. Journal of Artificial Intelligence Research. 65: 569–631.
Salaev U., Kuriyozov E., Gómez-Rodríguez C. 2022. SimRelUz: Similarity and relatedness scores as a semantic evaluation dataset for the Uzbek language. Proceedings of the 1st Annual Meeting of the ELRA/ISCA Special Interest Group on Under-Resourced Languages (SIGUL 2022). 199–206.
Sennrich R., Haddow B., Birch A. 2016. Neural machine translation of rare words with subword units. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (ACL 2016). Vol. 1. 1715–1725.
Weitzman Y., Hartmann M. 2025. Recent advancements and challenges of Turkic Central Asian language processing. Proceedings of the First Workshop on Language Models for Low-Resource Languages (LoResLM 2025). Abu Dhabi: ACL. arXiv:2407.05006.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 D.A. Akhmedjanova

This work is licensed under a Creative Commons Attribution 4.0 International License.