arXiv:2608.29120cs.CLcs.AI2026-09

让大模型学会根据声音判断说话人,提升多说话人场景下的推理能力。

HEAR Who Said What: Unlocking Speaker-Attributed Reasoning via Counterfactual Voice Grounding

论文配图:HEAR Who Said What: Unlocking Speaker-Attributed Reasoning via Counterfactual Voice Grounding
图 1 · 摘自论文原文
  • 构建分层评测基准HEAR,验证模型对说话人身份的准确判断能力。
  • 30B模型A2R在HEAR上表现优异,且能零样本泛化到多种下游任务。
  • 通过反事实语音数据训练,让模型更依赖声音特征而非语义内容。

语音语言模型(SLMs)在多说话人场景中应用日益广泛,但其对说话人身份的识别与推理能力尚不明确。为此,我们提出HEAR,一个概念分层的评测基准,包含从887段多样多对话音频中筛选的2.4K条人工验证样本。在HEAR上评估20个主流SLMs发现,它们普遍依赖语义先验而非实际声学线索。为此,我们引入A2R,一个基于反事实语音与说话人级困难负样本(CASH)数据集训练的30B模型,旨在引导模型优先关注声学特征。A2R在HEAR上表现良好,并展现出对多样化多说话人下游任务的零样本泛化能力,表明学习说话人归属可释放模型潜在的说话人感知推理能力。所有资源详见 https://attributetoreason.github.io/AttributeToReason/

原文摘要 · Abstract (English)

Speech Language Models (SLMs) are increasingly deployed in multi-speaker environments, yet their ability to attribute speech to the correct speaker and reason over speaker identities remains unclear. Hence, we introduce HEAR, a conceptually hierarchical benchmark diagnosing the foundational capabilities of speaker-attributed reasoning, comprising 2.4K human-verified samples from 887 diverse multi-party audio clips. Evaluating 20 leading SLMs on HEAR reveals they struggle with these foundational tasks, often relying on semantic priors rather than actual vocal cues. To address this, we present A2R, a 30B model optimized on Counterfactual Audio with Speaker-level Hard negatives (CASH), a dataset designed to guide the model to prioritize acoustic vocal cues over linguistic signals. A2R achieves strong performance on HEAR and exhibits zero-shot generalization to diverse multi-speaker downstream tasks, demonstrating that learned speaker attribution unlocks the model's latent capacity for speaker-aware reasoning. All resources are available at https://attributetoreason.github.io/AttributeToReason/

语音识别说话人识别大模型推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。