arXiv:2606.19157eess.AScs.CL2026-06中稿 · Interspeech 2026被引 1

评测语音大模型如何真正用好上下文,针对8种印度语言设计新基准

IndicContextEval: A Benchmark for Evaluating Context Utilisation in Audio Large Language Models Across 8 Indic Languages

  • 设计7级提示框架,逐步加入元数据、实体列表等上下文信息
  • 在56小时多语种数据上测试5个模型,发现上下文使用差异显著
  • 适合关注语音模型真实理解能力的研究者和开发者

语音大模型(AudioLLMs)支持基于文本提示(如领域描述或实体列表)的语音识别。然而,尚不清楚这些模型是否真正利用了上下文,还是仅依赖预训练中学习的参数化知识。现有基准无法回答此问题,因其评估条件固定且很少包含显式上下文输入。本文提出IndicContextEval,一个涵盖8种印度语言、23个专业领域的56小时多语种自然语音数据集,来自555名说话者。设计了7级提示框架,逐步引入元数据、自然语言描述、英文与本地文字的实体列表,以及含错误实体的对抗性提示。对5个模型的评估揭示了上下文利用行为的显著差异,凸显了对语音大模型上下文关联能力进行显式评估的必要性。

原文摘要 · Abstract (English)

AudioLLMs enable speech recognition conditioned on textual prompts such as domain descriptions or entity lists. However, it remains unclear whether these models genuinely utilise such context or rely on parametric knowledge learned during pretraining. Existing benchmarks cannot answer this question because they evaluate transcription under fixed prompting conditions and rarely include explicit contextual inputs. We introduce IndicContextEval, a 56-hour multilingual benchmark of natural speech from 555 speakers across 8 Indian languages and 23 professional domains. We design a 7-level prompting framework that progressively introduces contextual signals, including metadata, natural-language descriptions, entity lists in English and native script, and adversarial prompts with incorrect entities. Evaluating five models reveals substantial differences in context utilisation behaviour, highlighting the need for explicit evaluation of contextual grounding in AudioLLMs.

语音识别上下文理解多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。