arXiv:2510.19144cs.CL2025-10综述被引 1

系统梳理藏语AI研究现状,助力低资源语言技术发展

Tibetan Language and AI: A Comprehensive Survey of Resources, Methods and Challenges

  • 全面整理藏语文本与语音数据集及工具资源
  • 揭示数据稀缺、拼写不统一等核心瓶颈问题
  • 适合关注低资源语言与跨语言迁移的研究者

藏语是亚洲主要的低资源语言之一,具有独特的语言和文化特征,为人工智能研究带来挑战与机遇。尽管对代表性不足语言的AI系统开发兴趣日益增长,但由于缺乏可访问的数据资源、标准化基准和专用工具,藏语仍受到有限关注。本文系统综述了当前藏语AI研究的进展,涵盖文本与语音数据资源、自然语言处理任务、机器翻译、语音识别以及大模型(LLM)的最新发展。我们对现有数据集和工具进行分类,评估各任务所用方法并比较性能。同时指出持续存在的瓶颈,如数据稀疏性、拼写变体及缺乏统一评估指标。此外,讨论了跨语言迁移、多模态学习和社区驱动资源建设的潜力。本综述旨在为未来藏语AI研究提供基础参考,并鼓励建立包容且可持续的低资源语言AI生态。

原文摘要 · Abstract (English)

Tibetan, one of the major low-resource languages in Asia, presents unique linguistic and sociocultural characteristics that pose both challenges and opportunities for AI research. Despite increasing interest in developing AI systems for underrepresented languages, Tibetan has received limited attention due to a lack of accessible data resources, standardized benchmarks, and dedicated tools. This paper provides a comprehensive survey of the current state of Tibetan AI in the AI domain, covering textual and speech data resources, NLP tasks, machine translation, speech recognition, and recent developments in LLMs. We systematically categorize existing datasets and tools, evaluate methods used across different tasks, and compare performance where possible. We also identify persistent bottlenecks such as data sparsity, orthographic variation, and the lack of unified evaluation metrics. Additionally, we discuss the potential of cross-lingual transfer, multi-modal learning, and community-driven resource creation. This survey aims to serve as a foundational reference for future work on Tibetan AI research and encourages collaborative efforts to build an inclusive and sustainable AI ecosystem for low-resource languages.

藏语AI低资源语言自然语言处理数据资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。