专业语音识别评测基准揭示模型对上下文信息利用不足的问题。
PROFASR-BENCH: A Benchmark for Context-Conditioned ASR in High-Stakes Professional Speech
- 构建多领域专业对话数据集,支持带上下文的语音识别测试
- 发现现有模型在加入上下文后词错误率几乎不变,存在上下文利用缺口
- 提供可复现的评测平台,适合研究上下文融合策略的学者
专业场景下的自动语音识别面临现有评测基准未能充分反映的挑战:密集的专业术语、正式语体变化以及对关键实体错误近乎零容忍。我们提出ProfASR-Bench,一个面向金融、医疗、法律和技术等高风险应用领域的专业话语评估套件。每个样本包含自然语言提示(领域线索和/或说话人画像)与富含实体的目标语句,实现对上下文感知识别的可控测量。该语料库支持传统ASR指标、实体感知评分及按口音和性别分片的报告。在匹配的无上下文、画像、领域+画像、理想提示和对抗性提示条件下,测试代表性的Whisper(编码器-解码器ASR)和Qwen-Omni(音频语言模型),结果一致显示:轻量文本上下文对平均词错误率(WER)影响微乎其微,即使使用理想提示也未显著改善;对抗性提示亦不能稳定降低性能。我们称此为上下文利用缺口(CUG):当前系统名义上可提示,却严重低估可用的辅助信息。ProfASR-Bench提供标准化的上下文阶梯、实体与分片感知报告(含置信区间)以及可复现的测试环境,用于跨模型族融合策略的比较。
原文摘要 · Abstract (English)
Automatic Speech Recognition (ASR) in professional settings faces challenges that existing benchmarks underplay: dense domain terminology, formal register variation, and near-zero tolerance for critical entity errors. We present ProfASR-Bench, a professional-talk evaluation suite for high-stakes applications across finance, medicine, legal, and technology. Each example pairs a natural-language prompt (domain cue and/or speaker profile) with an entity-rich target utterance, enabling controlled measurement of context-conditioned recognition. The corpus supports conventional ASR metrics alongside entity-aware scores and slice-wise reporting by accent and gender. Using representative families Whisper (encoder-decoder ASR) and Qwen-Omni (audio language models) under matched no-context, profile, domain+profile, oracle, and adversarial conditions, we find a consistent pattern: lightweight textual context produces little to no change in average word error rate (WER), even with oracle prompts, and adversarial prompts do not reliably degrade performance. We term this the context-utilization gap (CUG): current systems are nominally promptable yet underuse readily available side information. ProfASR-Bench provides a standardized context ladder, entity- and slice-aware reporting with confidence intervals, and a reproducible testbed for comparing fusion strategies across model families. Dataset: https://huggingface.co/datasets/prdeepakbabu/ProfASR-Bench Code: https://github.com/prdeepakbabu/ProfASR-Bench
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。