让AI通过对话动态学习低资源语言,打破数据依赖瓶颈。
Towards Open-Ended Discovery for Low-Resource NLP
- AI与人类通过对话协作,动态发现新语言
- 结合模型不确定性与人类反馈信号,优化学习过程
- 适合关注人机共学与语言多样性保护的研究者
低资源语言的自然语言处理受限于文本语料匮乏、标准拼写缺失及可扩展的标注流程不足。尽管大语言模型提升了跨语言迁移能力,但其对海量预收集数据和集中式基础设施的依赖,使边缘群体难以受益。本文主张转向开放式的交互式语言发现范式,即AI通过持续对话而非静态数据集动态学习新语言。我们提出基于人机联合不确定性的框架,融合模型的认知不确定性与人类说话者的犹豫提示和信心信号,以指导互动、查询选择与记忆保留。该愿景强调人本化AI,推动从数据提取转向参与式、协同适应的学习过程,尊重并赋能社区,助力全球语言多样性的发现与保存。
原文摘要 · Abstract (English)
Natural Language Processing (NLP) for low-resource languages remains fundamentally constrained by the lack of textual corpora, standardized orthographies, and scalable annotation pipelines. While recent advances in large language models have improved cross-lingual transfer, they remain inaccessible to underrepresented communities due to their reliance on massive, pre-collected data and centralized infrastructure. In this position paper, we argue for a paradigm shift toward open-ended, interactive language discovery, where AI systems learn new languages dynamically through dialogue rather than static datasets. We contend that the future of language technology, particularly for low-resource and under-documented languages, must move beyond static data collection pipelines toward interactive, uncertainty-driven discovery, where learning emerges dynamically from human-machine collaboration instead of being limited to pre-existing datasets. We propose a framework grounded in joint human-machine uncertainty, combining epistemic uncertainty from the model with hesitation cues and confidence signals from human speakers to guide interaction, query selection, and memory retention. This paper is a call to action: we advocate a rethinking of how AI engages with human knowledge in under-documented languages, moving from extractive data collection toward participatory, co-adaptive learning processes that respect and empower communities while discovering and preserving the world's linguistic diversity. This vision aligns with principles of human-centered AI, emphasizing interactive, cooperative model building between AI systems and speakers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。