arXiv:2411.13409cs.CLcs.AI2024-11中稿 · IEEE conference IS…

用大模型技术统一濒危的巴尔蒂语方言,促进跨区域语言融合。

Unification of Balti and trans-border sister dialects in the essence of LLMs and AI Technology

  • 利用大语言模型分析不同地区的巴尔蒂语方言差异。
  • 通过跨文化数据整合,构建统一的词汇与语音体系。
  • 适合语言保护、AI多语种研究者参考。

巴尔蒂语属于汉藏语系藏缅语族,分布于印度、中国、巴基斯坦、尼泊尔、西藏、缅甸和不丹等地,受当地文化影响产生多种方言。在全球化与人工智能技术快速发展的背景下,理解方言多样性并推动其统一,有助于揭示共同语言根基,缩小因地理、社会政治及宗教差异造成的沟通鸿沟。本文探讨大语言模型(LLMs)如何在分析、记录与标准化濒危巴尔蒂语方面发挥作用,基于已有各地区方言研究工作,提出以人工智能技术为支撑的语言统一路径。

原文摘要 · Abstract (English)

The language called Balti belongs to the Sino-Tibetan, specifically the Tibeto-Burman language family. It is understood with variations, across populations in India, China, Pakistan, Nepal, Tibet, Burma, and Bhutan, influenced by local cultures and producing various dialects. Considering the diverse cultural, socio-political, religious, and geographical impacts, it is important to step forward unifying the dialects, the basis of common root, lexica, and phonological perspectives, is vital. In the era of globalization and the increasingly frequent developments in AI technology, understanding the diversity and the efforts of dialect unification is important to understanding commonalities and shortening the gaps impacted by unavoidable circumstances. This article analyzes and examines how artificial intelligence AI in the essence of Large Language Models LLMs, can assist in analyzing, documenting, and standardizing the endangered Balti Language, based on the efforts made in different dialects so far.

语言保护大模型多语种

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。