用少量语音样本让模型像人一样适应不同口音和说话人,提升语音识别准确率。
In-Context Learning Boosts Speech Recognition via Human-like Adaptation to Speakers and Language Varieties
- 通过交错的语音-文本提示实现推理时上下文学习,仅需12段约50秒样本
- 跨多种英语语料平均降低相对19.7%的词错误率(绝对1.2个百分点)
- 对低资源口音效果更显著,适合需要鲁棒语音识别的场景
人类听众可通过暴露快速适应陌生说话人和语言变体,这一适应能力是否能延伸至先进的语音语言模型?我们提出一种可扩展框架,使Phi-4 Multimodal在推理时支持上下文学习(ICL),采用交错的任务提示与音频-文本对。结果表明,仅需12个示例语句(约50秒)即可在多样化的英语语料上平均降低相对19.7%(绝对1.2个百分点)的词错误率。该改进在低资源语言变体中最为显著,且当上下文与目标说话人匹配、示例数量更多时效果更佳;但随着上下文长度增加,收益呈现边际递减。总体而言,我们的新式ICL适配方案(1)展现出与人类听众相似的表现特征,(2)在不同说话人和语言背景中一致提升了自动语音识别(ASR)的鲁棒性。尽管适应普遍有效,某些语言变体仍存在显著差距,揭示当前模型在人类灵活性方面仍有不足。我们已在GitHub发布提示模板与代码。
原文摘要 · Abstract (English)
Human listeners readily adjust to unfamiliar speakers and language varieties through exposure, but do these adaptation benefits extend to state-of-the-art spoken language models? We introduce a scalable framework that allows for in-context learning (ICL) in Phi-4 Multimodal using interleaved task prompts and audio-text pairs, and find that as few as 12 example utterances (~50 seconds) at inference time reduce word error rates by a relative 19.7% (1.2 pp.) on average across diverse English corpora. These improvements are most pronounced in low-resource varieties, when the context and target speaker match, and when more examples are provided--though scaling our procedure yields diminishing marginal returns to context length. Overall, we find that our novel ICL adaptation scheme (1) reveals a similar performance profile to human listeners, and (2) demonstrates consistent improvements to automatic speech recognition (ASR) robustness across diverse speakers and language backgrounds. While adaptation succeeds broadly, significant gaps remain for certain varieties, revealing where current models still fall short of human flexibility. We release our prompts and code on GitHub.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。