arXiv:2510.22492cs.CL2025-10被引 1

发现多语言语音识别中音素使用存在声学饱和阈值,数据量不是决定因素。

The Limits of Data Scaling: Sub-token Utilization and Acoustic Saturation in Multilingual ASR

  • 通过追踪解码过程中的子词发现,分析模型在49种语言中的使用模式。
  • 子词发现率呈指数饱和,约70%的子词在前30分钟内被激活,后续增长极慢。
  • 语言类型(如拉丁字母)比非拉丁文字更高效,提示应更关注语言结构而非数据量。

我们研究了多语言语音识别模型在49种语言中解码行为的子词使用规律。通过记录解码候选子词并追踪其累积发现过程,发现已发现子词总数与预训练时长无关,表明数据差异对模型词汇多样性影响有限。子词发现速率在各语言中均呈现一致的指数饱和趋势,存在一个收敛阈值,称为声学饱和时间(AST)。进一步分析显示,子词频次分布符合齐夫-曼德尔布罗特定律,平均子词长度与资源水平正相关。拉丁字母语言在各项指标上表现更优,而斯拉夫、中日韩及闪米特系文字表现较差。结果表明,多语言语音识别中的子词利用受限于语言的统计、类型和书写系统结构,而非训练数据规模,为更公平的数据构建和跨语言评估提供实证依据。

原文摘要 · Abstract (English)

How much audio is needed to fully observe a multilingual ASR model's learned sub-token inventory across languages, and does data disparity in multilingual pre-training affect how these tokens are utilized during inference? We address this question by analyzing Whisper's decoding behavior during inference across 49 languages. By logging decoding candidate sub-tokens and tracking their cumulative discovery over time, we study the utilization pattern of the model's sub-token space. Results show that the total number of discovered tokens remains largely independent of a language's pre-training hours, indicating that data disparity does not strongly influence lexical diversity in the model's hypothesis space. Sub-token discovery rates follow a consistent exponential saturation pattern across languages, suggesting a stable time window after which additional audio yields minimal new sub-token activation. We refer to this convergence threshold as acoustic saturation time (AST). Further analyses of rank-frequency distributions reveal Zipf-like patterns better modeled by a Zipf-Mandelbrot law, and mean sub-token length shows a positive correlation with resource level. Additionally, those metrics show more favorable patterns for languages in the Latin script than those in scripts such as Cyrillic, CJK, and Semitic. Together, our study suggests that sub-token utilization during multilingual ASR inference is constrained more by the statistical, typological, and orthographic structure of the speech than by training data scale, providing an empirical basis for more equitable corpus construction and cross-lingual evaluation.

语音识别多语言子词声学饱和

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。