让语言识别更包容:用环境线索提升小语种识别率
How Should We Model the Probability of a Language?
- 将语言识别视为路由问题,而非孤立文本分类
- 利用环境线索提升低资源语言的识别可行性
- 适合关注多语言公平性与系统设计的研究者
全球7000多种语言中,商业级语言识别(LID)系统仅能可靠识别少数几百种书面语言。研究级系统在特定条件下可扩展覆盖范围,但大多数语言仍缺乏有效支持。本文认为,这种局面主要源于对LID的错误定位——将其视为脱离上下文的文本分类任务,从而忽视了先验概率估计的核心作用,并受制于推崇全球统一先验模型的制度激励。我们主张,要提升长尾语言的覆盖,必须重新将LID理解为一种路由问题,并建立可解释的方法,融入使语言在当地语境中合理的环境线索。
原文摘要 · Abstract (English)
Of the over 7,000 languages spoken in the world, commercial language identification (LID) systems only reliably identify a few hundred in written form. Research-grade systems extend this coverage under certain circumstances, but for most languages coverage remains patchy or nonexistent. This position paper argues that this situation is largely self-imposed. In particular, it arises from a persistent framing of LID as decontextualized text classification, which obscures the central role of prior probability estimation and is reinforced by institutional incentives that favor global, fixed-prior models. We argue that improving coverage for tail languages requires rethinking LID as a routing problem and developing principled ways to incorporate environmental cues that make languages locally plausible.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。