arXiv:2505.18159cs.CLcs.LG2025-05被引 1

首个针对濒危语言康麦奇语的计算研究,用少量数据提升大模型识别准确率。

Advancing Uto-Aztecan Language Technologies: A Case Study on the Endangered Comanche Language

  • 构建412条短语的手工数据集,结合合成数据增强模型训练。
  • 仅用5个示例的少样本提示,使GPT-4o实现接近100%的识别准确率。
  • 强调社区参与与文化敏感性,为濒危语言提供可复制的数字保护路径。

濒危语言在自然语言处理中的数字排斥仍是重大挑战,制约语言学研究与复兴工作。本研究首次对濒临灭绝的乌托-阿兹特克语系康麦奇语开展计算探索,展示了低成本、社区参与型NLP干预如何助力语言保存。我们构建了一个包含412条短语的手工标注数据集,设计了一套合成数据生成流程,并对GPT-4o与GPT-4o-mini进行了语言识别的实证评估。实验表明,尽管大模型在零样本设置下表现不佳,但通过少样本提示(仅需5个示例)显著提升性能,达到近乎完美的准确率。研究凸显了针对性NLP方法在低资源场景下的潜力,强调可见性是实现包容的第一步。通过为康麦奇语建立NLP基础,本文倡导以可访问性、文化敏感性和社区协作为核心的计算路径。

原文摘要 · Abstract (English)

The digital exclusion of endangered languages remains a critical challenge in NLP, limiting both linguistic research and revitalization efforts. This study introduces the first computational investigation of Comanche, an Uto-Aztecan language on the verge of extinction, demonstrating how minimal-cost, community-informed NLP interventions can support language preservation. We present a manually curated dataset of 412 phrases, a synthetic data generation pipeline, and an empirical evaluation of GPT-4o and GPT-4o-mini for language identification. Our experiments reveal that while LLMs struggle with Comanche in zero-shot settings, few-shot prompting significantly improves performance, achieving near-perfect accuracy with just five examples. Our findings highlight the potential of targeted NLP methodologies in low-resource contexts and emphasize that visibility is the first step toward inclusion. By establishing a foundation for Comanche in NLP, we advocate for computational approaches that prioritize accessibility, cultural sensitivity, and community engagement.

濒危语言少样本学习社区参与语言复兴

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。