arXiv:2603.19468cs.SDeess.AS2026-03被引 1

让语音模型推理时标注时间戳,提升回答准确性与可信度

Listen First, Then Answer: Timestamp-Grounded Speech Reasoning

  • 用强化学习让模型在推理中显式标注音频时间戳
  • 在四个语音数据集上性能优于零样本与无时间戳微调
  • 增强模型对音频片段的注意力和推理一致性,适合需要可解释性的场景

大型音频-语言模型(LALMs)能生成预测的推理链,但这些推理是否真正基于输入音频仍不明确。本文提出一种基于强化学习的策略,通过显式的时间戳标注,将模型推理结果锚定到音频信号的相关片段。分析显示,时间戳锚定使模型在推理过程中更强烈地关注音频标记。在四个语音基准数据集上的实验表明,该方法在性能上优于零样本推理和未使用时间戳锚定的微调。此外,锚定机制还增强了区域探索、听觉验证和一致性等理想推理行为,凸显了锚定机制对真实多模态推理的重要性。

原文摘要 · Abstract (English)

Large audio-language models (LALMs) can generate reasoning chains for their predictions, but it remains unclear whether these reasoning chains remain grounded in the input audio. In this paper, we propose an RL-based strategy that grounds the reasoning outputs of LALMs with explicit timestamp annotations referring to relevant segments of the audio signal. Our analysis shows that timestamp grounding leads the model to attend more strongly to audio tokens during reasoning generation. Experiments on four speech-based benchmark datasets demonstrate that our approach improves performance compared to both zero-shot reasoning and fine-tuning without timestamp grounding. Additionally, grounding amplifies desirable reasoning behaviors, such as region exploration, audiology verification, and consistency, underscoring the importance of grounding mechanisms for faithful multimodal reasoning.

语音理解多模态推理时间戳标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。