Digital Transformation of the Georgian Language: The Role of Artificial Intelligence and Language Technologies
Downloads
The rapid development of digital technologies has significantly changed the communicative, educational and information environment of modern society. The integration of artificial intelligence and language technologies has acquired particular importance, creating new opportunities for both global and less common languages. The Georgian language, as a language with a unique script and centuries-old cultural heritage, faces significant challenges in the modern digital era. On the one hand, its full integration into modern information technologies is necessary, and on the other hand, the protection of linguistic identity and norms is necessary. The article discusses modern trends in the digital transformation of the Georgian language, the role of artificial intelligence in the development of language resources, the capabilities of natural language processing technologies, automatic translation systems, speech recognition and synthesis platforms, the impact of automatic spelling and grammar checking corpora, as well as large language models on the digital development of the Georgian language. Special attention is paid to Georgian language corpora, the need to develop lexical resources and digital infrastructure.
The goal of the study is to analyze the main directions of the digital transformation of the Georgian language, assess the existing challenges, and determine strategic steps that will contribute to increasing the competitiveness of the Georgian language in the global digital space. The findings indicate that the successful digital transformation of the Georgian language largely depends on the development of comprehensive language resources, open data infrastructure, and sustainable collaboration between academic institutions, government agencies, and the technology sector. In addition, the integration of the Georgian language into
modern artificial intelligence systems requires a focused effort in language modeling and increased investment in computational linguistics research. The article states that the digital existence of the Georgian language is not only a technological issue, but also a matter of cultural sustainability and national strategic development in the era of
artificial intelligence.
Downloads
„ქართული ენის ეროვნული კორპუსი“, ქეეკ, https://surl.li/cjupvp
მეცნიერებათა ეროვნული აკადემია „მაცნე“ 2024, https://doi.org/10.48550/arXiv.2405.00710
https://arxiv.org/pdf/2405.00710
სამოქალაქო ინტეგრაციის და ეროვნებათშორისი ურთიერთობების ცენტრი, „ენობრივი პოლიტიკა საქართველოში: სიტუაციის ანალიზი და კვლევის შედეგები“, 2023, https://surl.li/ediraz
ტაბიძე, მ. „ენობრივი პოლიტიკა“, ეს გვერდი ბოლოს განახლდა 23:18, 30 იანვარი 2024. https://surl.li/wgmsco
ქართული ენის კორპუსი, ილიას სახელმწიფო უნივერსიტეტის ენათმეცნიერების ინსტიტუტი, https://surl.li/furlsn
ქეთევან მჭედლიშვილი, „შედარებითი კორპუსები: შედგენის მეთოდოლოგია დაგამოყენების სფეროები“, კადმოსი. ჰუმანიტარულ კვლევათა ჟურნალი, 2024: 237-253, https://doi.org/10.32859/kadmos/16/237-253
Beso Mikaberidze, Teimuraz Saghinadze, Guram Mikaberidze, Raphael Kalandadze, Konstantine Pkhakadze, Josef van Genabith, Simon Ostermann, Lonneke van der Plas, Philipp Müller, „A Comparison of Different Tokenization Methods for the Georgian Language“, Proceedings of the 7th International Conference on Natural Language and Speech Processing (ICNLSP 2024), https://surl.li/tgrvky
Beso Mikaberidze, Teimuraz Saghinadze, Guram Mikaberidze, Raphael Kalandadze, Konstantine Pkhakadze, Josef van Genabith, Simon Ostermann, Lonneke van der Plas, Philipp Müller, „A Comparison of Different Tokenization Methods for the Georgian Language“, Proceedings of the 7th International Conference on Natural Language and Speech Processing (ICNLSP 2024)
Computer Science > Computation and Language, https://arxiv.org/abs/2303.08774
Computer Science > Computation and Language, OpenAI. (2023). https://arxiv.org/abs/2303.08774
Daniel Gallagher, Gerhard Heyer, Computer Science-Computation and Language, „Targeted Syntactic Evaluation of Language Models on Georgian Case Alignment“, https://doi.org/10.48550/arXiv.2602.10661
David Ifeoluwa Adelani, Graham Neubig, Sebastian Ruder, Shruti Rijhwani, Michael Beukman, Chester Palen-Michel, Constantine Lignos, Jesujoba O. Alabi, Shamsuddeen H. Muhammad, Peter Nabende, „MasakhaNER 2.0: Africa-centric Transfer Learning for Named Entity Recognition,“ EMNLP 2022 . https://surl.li/pdiwjx
David Ifeoluwa Adelani, Md Mahfuz Ibn Alam, Antonios Anastasopoulos, Akshita Bhagia, Marta R. Costa-jussà, Jesse Dodge, Fahim Faisal, Christian Federmann, Natalia Fedorova, Francisco Guzmán, Sergey Koshelev, Jean Maillard, Vukosi Marivate, Jonathan Mbuya, Alexandre Mourachko, Safiyyah Saleem, Holger Schwenk, Guillaume Wenzek, „Findings of the WMT’22 Shared Task on Large-Scale Machine Translation Evaluation for African Languages“, Volume: Proceedings of the Seventh Conference on Machine Translation (WMT), December, 2022. https://surl.li/yxlqqe
Davit Melikidze, Alexander Gamkrelidze, „Homonym Sense Disambiguation in the Georgian Language“,
Emily M. Bender, Batya Friedman, „Data Statements for Natural Language Processing: Toward Mitigating System Bias and Enabling Better Science“, Transactions of the Association for Computational Linguistics, vol. 6, pp. 587–604, 2018. Action Editor: Yuji Matsumoto. Submission batch: 5/2018; Revision batch: 8/2018; Published 12/2018. https://aclanthology.org/Q18-1041.pdf
European Commission „Language Technologies and Multilingualism“, https://surl.li/wkqvks, Last update, 23 June 2026
Fajri Koto, Afshin Rahimi, Jey Han Lau, Timothy Baldwin, „IndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP, November 2020,
Internationalization (i18n), Making the World Wide Web worldwide!, https://www.w3.org/International/
Irina Lobzhanidze, Erekle Magradze, Svetlana Berikashvili, Anzor Gozalishvili, Tamar Jalaghonia, „Building a Universal Dependencies Treebank for Georgian“, Publisher: Association for Computational Linguistics, Volume: Proceedings of the 22nd Workshop on Treebanks and Linguistic Theories (TLT 2024), https://aclanthology.org/2024.tlt-1.5/
Jerome Aondongu Achir, Matthew T Ogedengbe, Joseph Sarwuan, „Preservation of Low Resource Languages through Natural Language Processing: Challenges, Opportunities, and the Case of Tiv“, Jurnal of Basics and Applied Sciences Researcher (JOBASR), Volume 4(2), 2026 DOI: https://dx.doi.org/10.4314/jobasr.v4i2.21
Jurafsky, D. and Martin, J. H., „Speech and Language Processing (3rd ed., draft).“ Stanford University. (2023), https://surli.cc/gimopn
Kamarauli, M. (2024). Enhancement possibilities for the Georgian National Corpus . Caucasus Journal of Social Sciences, 17(1), 142–165. https://doi.org/10.62343/cjss.2024.245
Manning, S. D. and Schutz, H „Fundamentals of Statistical Natural Language Processing“, (1999). https://surl.li/vuvkns
Mariam Kamarauli, „Enhancement possibilities for the Georgian National Corpus“, 2024, https://surl.li/qywojh https://doi.org/10.62343/cjss.2024.245,
Nicolas Stefanovitch, Jakub Piskorski, Sopho Kharazi, „Resources and Experiments on Sentiment Classification for Georgian“, Proceedings of the Thirteenth International Conference on Language Resources and Evaluation (LREC 2022), https://doi.org/10.63317/4ky6so2mrss6
OECD, 2024, „Policies, data and analysis for trustworthy artificial intelligence“, https://oecd.ai/en/
Paul Meurer „A Computational Grammar for Georgian“, Conference: Logic, Language, and Computation, 7th International Tbilisi Symposium on Logic, Language, and Computation, TbiLLC 2007, Tbilisi, Georgia, October 1-5, 2007. Revised Selected Papers, https://doi.org/10.1007/978-3-642-00665-4_1
Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, „The State and Fate of Linguistic Diversity and Inclusion in the NLP World“, Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 6282–6293 July 5- 10, 2020. https://surl.li/mhhkhu
Proceedings of the 13th Conference on Language Resources and Assessment, European Language Resources Association LREC (2022–2024). https://lrec-conf.org
Rishi Bommasani, Drew A. Hudson, Ehsan Adeli, „On the Opportunities and Risks of Foundation Models„, August 2021, DOI: 10.48550/arXiv.2108.07258 https://surl.li/lmagxm
Rodrigo Agerri, Iñaki San Vicente, Jon Ander Campos, Ander Barrena, Xabier Saralegi, Aitor Soroa, Eneko Agirre, „Give your Text Representation Models some Love: the Case for Basque“, Publisher: European Language Resources Association, Proceedings of the Twelfth Language Resources and Evaluation Conference, 2020. https://aclanthology.org/2020.lrec-1.588/
Steven Bird, Ewan Klein, Edward Loper, „Natural Language Processing with Python“, Publisher: O'Reilly, ISBN: 978-0-596-51649-9January 2009, https://surl.li/ysucjv
Tom. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, „Language Models are Few-Shot Learners“, 34th Conference on Neural Information Processing Systems (NeurIPS 2020), Vancouver, https://surl.li/ecijgw
UNESCO, 2021, „Recommendation on the Ethics of Artificial Intelligence“ https://surl.li/givvwl
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. „Attention Is All You Need“ (2017). https://arxiv.org/abs/1706.03762
Copyright (c) 2026 Georgian Scientists

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.

