构建真实场景下的定制词汇语音识别基准,推动工业级语音转写落地。
Contextual Earnings-22: A Speech Recognition Benchmark with Custom Vocabulary in the Wild

- 基于Earnings-22构建含真实定制词汇的开放数据集
- 关键词提示与增强方法在大规模下准确率显著提升
- 适合研究工业场景语音识别的团队和开发者
学术语音识别基准的准确率已趋于饱和,而工业应用和高风险领域实践表明仍有进步空间。我们提出核心差异在于上下文条件:学术数据以常见通用词汇为主,易于识别;而真实场景中的罕见、定制化词汇对转录可用性影响更大。尽管已有上下文语音识别进展,但缺乏标准化评测基准。本文引入Contextual Earnings-22,基于Earnings-22构建包含真实定制词汇语境的开放数据集,以促进研究并揭示潜在进展。针对主流两种方法(关键词提示与关键词增强),设立六组强基线。实验显示,二者在从概念验证扩展到大规模系统时均达到可比且显著提升的准确率。
原文摘要 · Abstract (English)
The accuracy frontier of speech-to-text systems has plateaued on academic benchmarks.1 In contrast, industrial benchmarks and adoption in high-stakes domains suggest otherwise. We hypothesize that the primary difference between the two is contextual conditioning: Academic benchmarks are dominated by frequently encountered general vocabulary that is relatively easy to recognize compared with rare and context-defined custom vocabulary that has disproportionate impact on the usability of speech transcripts. Despite progress on contextual speech-to-text, there is no standardized benchmark. We introduce Contextual Earnings-22, an open dataset built upon Earnings-22, with realistic custom vocabulary contexts to foster research and reveal latent progress. We set six strong baselines for two dominant approaches: keyword prompting and keyword boosting. Experiments show both reach comparable and significantly improved accuracy when scaled from proof-of-concept to large-scale systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。