厘清低资源语言的定义困境,揭示其多维成因
The Zeno's Paradox of `Low-Resource' Languages
- 分析150篇论文,发现低资源语言无统一标准
- 提出语言资源匮乏由多重因素共同作用所致
- 呼吁研究者明确定义术语,推动可比性研究
自然语言处理领域常将语言分为高资源与低资源两类,但何为‘低资源语言’缺乏共识。为探究该术语在论文中的实际使用方式,我们对来自ACL Anthology及主流语音处理会议中提及‘low-resource’的150篇论文进行了定性分析。结果表明,语言资源匮乏是多个相互关联维度共同作用的结果,包括语料规模、标注数据量、工具支持程度和社区活跃度等。这些维度的复杂交织使得难以针对单个语言建立一致的评估基准,也阻碍了跨语言进展的追踪。本文旨在(1)促使研究者在使用该术语时给出明确界定;(2)提供一个系统性的分析框架,用于判断一种语言是否属于低资源范畴。
原文摘要 · Abstract (English)
The disparity in the languages commonly studied in Natural Language Processing (NLP) is typically reflected by referring to languages as low vs high-resourced. However, there is limited consensus on what exactly qualifies as a `low-resource language.' To understand how NLP papers define and study `low resource' languages, we qualitatively analyzed 150 papers from the ACL Anthology and popular speech-processing conferences that mention the keyword `low-resource.' Based on our analysis, we show how several interacting axes contribute to `low-resourcedness' of a language and why that makes it difficult to track progress for each individual language. We hope our work (1) elicits explicit definitions of the terminology when it is used in papers and (2) provides grounding for the different axes to consider when connoting a language as low-resource.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。