arXiv:2606.04474cs.CLeess.AS2026-06

语音大模型推理时因实体绑定失败导致逻辑任务表现差,用显式绑定可显著提升性能。

Entity Binding Failures in Speech LLM Reasoning: Diagnosis and Chain-of-Thought Intervention

  • 引入实体感知的思维链,推理前显式绑定实体与属性
  • 在逻辑任务上使准确率提升24.4个百分点,即使语音识别出错也有效
  • 揭示问题本质是推理激发不足,而非模型能力缺失

语音大语言模型(SLLMs)在复杂推理任务上的表现逊于文本模型。我们发现这一差距并非普遍认知缺陷。评估两种架构不同的SLLM后发现,语音转文本(S2T)在空间、句法和事实类任务上表现与文本转文本(T2T)相当或更优。但在需要实体追踪的逻辑任务中,S2T准确率跌至随机水平。我们诊断此为实体绑定失败:连续语音特征在隐式推理过程中模糊了实体与属性的精确关联。为验证该诊断,我们提出轻量级推理时干预方法——实体感知思维链(EA-CoT),强制SLLM在推理前枚举实体并将其与主张绑定。EA-CoT成功弥合差距,即使语音名称识别错误,仍实现最高达24.4个百分点的准确率提升。消融实验确认增益源于显式语义绑定,将差距重新定义为激发失败,而非能力缺失。

原文摘要 · Abstract (English)

Speech Large Language Models (SLLMs) underperform their text counterparts on complex reasoning. We reveal that this gap is not a uniform cognitive deficit. Evaluating two architecturally diverse SLLMs, we show speech-to-text (S2T) matches or exceeds text-to-text (T2T) on spatial, syntactic, and factual tasks. Yet on logical tasks requiring entity tracking, S2T accuracy collapses to chance. We diagnose this as an entity binding failure: continuous speech features blur precise entity-property associations during implicit reasoning. To validate this diagnosis, we introduce Entity-Aware Chain-of-Thought (EA-CoT), a lightweight inference-time intervention forcing SLLMs to enumerate entities and bind them to claims before reasoning. EA-CoT bridges the gap, even when spoken names are misrecognized, yielding up to a 24.4 percentage-point accuracy gain. Ablations confirm the gains stem from explicit semantic binding, reframing the gap as an elicitation failure rather than a missing capability.

语音大模型推理增强实体绑定

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。