arXiv:2601.09413cs.SDcs.AI2026-01ACL被引 3

让语音模型学会自我判断该信自己还是外部感知,提升识别与音频推理可靠性。

Speech-Hands: A Self-Reflection Voice Agentic Approach to Speech Recognition and Audio Reasoning with Omni Perception

  • 引入可学习的自我反思机制,决定何时依赖自身判断、何时求助外部音频感知。
  • 在7个基准上比强基线降低12.1%的词错误率,音频问答准确率达77.37%。
  • 适用于复杂多选音频推理任务,适合追求鲁棒性语音智能的研究与应用。

我们提出一种语音代理框架,学习一项关键的全模态理解能力:判断何时信任自身,何时求助外部音频感知。研究受一个反直觉发现启发:将全模态模型同时微调用于语音识别和外部声音理解任务,常因噪声假设导致性能下降。为此,我们构建Speech-Hands框架,将问题重构为显式的自我反思决策。这一可学习的反思机制有效防止模型被错误的外部候选误导。实验表明,该代理行为机制能自然泛化至复杂的多选音频推理任务。在OpenASR排行榜上,Speech-Hands在七个基准上持续优于强基线,词错误率降低12.1%。模型在音频问答任务中达到77.37%准确率和高F1值,展现出在多样音频问答数据集上的鲁棒泛化与可靠性。通过统一感知与决策,本工作为构建更可靠、更坚韧的音频智能提供了实用路径。

原文摘要 · Abstract (English)

We introduce a voice-agentic framework that learns one critical omni-understanding skill: knowing when to trust itself versus when to consult external audio perception. Our work is motivated by a crucial yet counterintuitive finding: naively fine-tuning an omni-model on both speech recognition and external sound understanding tasks often degrades performance, as the model can be easily misled by noisy hypotheses. To address this, our framework, Speech-Hands, recasts the problem as an explicit self-reflection decision. This learnable reflection primitive proves effective in preventing the model from being derailed by flawed external candidates. We show that this agentic action mechanism generalizes naturally from speech recognition to complex, multiple-choice audio reasoning. Across the OpenASR leaderboard, Speech-Hands consistently outperforms strong baselines by 12.1% WER on seven benchmarks. The model also achieves 77.37% accuracy and high F1 on audio QA decisions, showing robust generalization and reliability across diverse audio question answering datasets. By unifying perception and decision-making, our work offers a practical path toward more reliable and resilient audio intelligence.

语音识别音频推理自反思代理系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。