arXiv:2501.11496cs.CLcs.AI2025-01被引 6

用AI助力濒危语言保护,关键在社区主导与伦理保障。

Generative AI and Large Language Models in Language Preservation: Opportunities and Challenges

  • 构建社区驱动的评估框架,平衡AI能力与语言风险。
  • 毛利语语音识别达92%准确率,但数据主权与模型偏见仍存挑战。
  • 适合语言保护者、研究者与政策制定者参考实践。

全球语言濒危危机迎来技术转折点,生成式人工智能(GenAI)与大语言模型(LLMs)为语料自动构建、转录、翻译和教学带来新可能。然而,碎片化实践与缺乏方法论,使数据稀缺、文化误用及伦理失误风险加剧。本文提出新型分析框架,系统评估GenAI应用与语言需求的匹配度,将社区治理与伦理防护置于核心。以毛利语复兴为例,该框架揭示社区主导的自动语音识别已实现92%准确率,同时凸显数字档案与教育工具中的数据主权缺失与模型偏见问题。研究强调,唯有基于社区中心的数据治理、持续评估与透明风险管理,才能让GenAI真正推动语言保护。此框架为研究者、语言社区与政策制定者提供不可或缺的伦理实践工具。

原文摘要 · Abstract (English)

The global crisis of language endangerment meets a technological turning point as Generative AI (GenAI) and Large Language Models (LLMs) unlock new frontiers in automating corpus creation, transcription, translation, and tutoring. However, this promise is imperiled by fragmented practices and the critical lack of a methodology to navigate the fraught balance between LLM capabilities and the profound risks of data scarcity, cultural misappropriation, and ethical missteps. This paper introduces a novel analytical framework that systematically evaluates GenAI applications against language-specific needs, embedding community governance and ethical safeguards as foundational pillars. We demonstrate its efficacy through the Te Reo Māori revitalization, where it illuminates successes, such as community-led Automatic Speech Recognition achieving 92% accuracy, while critically surfacing persistent challenges in data sovereignty and model bias for digital archives and educational tools. Our findings underscore that GenAI can indeed revolutionize language preservation, but only when interventions are rigorously anchored in community-centric data stewardship, continuous evaluation, and transparent risk management. Ultimately, this framework provides an indispensable toolkit for researchers, language communities, and policymakers, aiming to catalyze the ethical and high-impact deployment of LLMs to safeguard the world's linguistic heritage.

语言保护生成式AI社区治理大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。