arXiv:2602.08951cs.CL2026-02中稿 · Vardial 2026

让语言识别更包容:用环境线索提升小语种识别率

How Should We Model the Probability of a Language?

  • 将语言识别视为路由问题,而非孤立文本分类
  • 利用环境线索提升低资源语言的识别可行性
  • 适合关注多语言公平性与系统设计的研究者

全球7000多种语言中,商业级语言识别(LID)系统仅能可靠识别少数几百种书面语言。研究级系统在特定条件下可扩展覆盖范围,但大多数语言仍缺乏有效支持。本文认为,这种局面主要源于对LID的错误定位——将其视为脱离上下文的文本分类任务,从而忽视了先验概率估计的核心作用,并受制于推崇全球统一先验模型的制度激励。我们主张,要提升长尾语言的覆盖,必须重新将LID理解为一种路由问题,并建立可解释的方法,融入使语言在当地语境中合理的环境线索。

原文摘要 · Abstract (English)

Of the over 7,000 languages spoken in the world, commercial language identification (LID) systems only reliably identify a few hundred in written form. Research-grade systems extend this coverage under certain circumstances, but for most languages coverage remains patchy or nonexistent. This position paper argues that this situation is largely self-imposed. In particular, it arises from a persistent framing of LID as decontextualized text classification, which obscures the central role of prior probability estimation and is reinforced by institutional incentives that favor global, fixed-prior models. We argue that improving coverage for tail languages requires rethinking LID as a routing problem and developing principled ways to incorporate environmental cues that make languages locally plausible.

语言识别小语种先验建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。