arXiv:2606.11219cs.CLcs.AI2026-06ACL

评测语音模型在不同口音和场景下的语义推理能力

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents

论文配图:Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents
图 1 · 摘自论文原文
  • 构建五项语音语义推理任务,评估模型对口音与领域变化的适应性
  • 发现现有模型在口音变化下推理结果不稳定,易产生错误推断
  • 适合关注语音模型公平性与鲁棒性的研究者和开发者

语音语言模型(ALMs)在语音理解中应用日益广泛,但其在转录、文本到音频检索、字幕生成和问答之外的语义推理能力仍缺乏充分评估。特别是口音差异、领域迁移和语义过度推断的影响尚不明确。本文在五个语义与副语言推理任务上评估了音频语言模型:蕴含关系、一致性、合理性、口音漂移和口音约束。这些任务综合考察模型以语音为首要证据进行推理的能力,包括判断文本假设是否可由语音推出、反驳或无法确定;陈述是否与语音内容一致或冲突;论断在语境中是否合理;以及模型预测在口音变化下是否保持稳定或适当约束。结果揭示了当前音频推理评估的重大局限,并为更稳健、更公平的ALM设计与评估提供指导。

原文摘要 · Abstract (English)

Audio language models (ALMs) are increasingly used for speech-based understanding, yet their ability to perform semantic reasoning beyond transcription, Text-to-Audio Retrieval, Captioning, and Question-Answering accuracy remains insufficiently benchmarked. In particular, the effects of accent variation, domain shift, and semantic over-inference on audio reasoning are poorly understood. We evaluate audio language models across five semantic and paralinguistic reasoning tasks: entailment, consistency, plausibility, accent drift, and accent restraint. Collectively, these tasks assess a model's ability to reason over spoken audio as the primary evidence source, including whether a textual hypothesis can be inferred, contradicted, or left undetermined by the audio, whether statements align or conflict with spoken content, whether claims are plausible given the discourse, and whether model predictions remain stable or appropriately constrained across accent variation. These findings highlight critical limitations in current audio reasoning evaluations and hope to provide guidance for more robust and equitable ALM design and assessment

语音理解语义推理口音差异模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。