通过分步感知与推理,提升多说话人音频理解的准确性。
Listen, Pause, and Reason: Toward Perception-Grounded Hybrid Reasoning for Audio Understanding

- 分阶段训练:先感知声音属性,再优化推理过程。
- 在多个基准上超越基础模型,接近大规模模型表现。
- 适合需要精准音频分析的语音识别、会议记录等场景。
近期大型音频语言模型在音频理解方面表现出色,但常出现感知错误,且缺乏对结构化听觉场景的感知基础时难以实现可靠推理。受听觉场景分析启发,我们提出感知感知问答(PAQA)数据集,采用层次解耦策略分离语音与环境音,区分多位说话人,为训练提供明确的感知推理依据。在此基础上,我们提出两阶段混合感知-推理框架HyPeR:第一阶段在PAQA上微调模型以感知复杂音频中的声学属性;第二阶段利用GRPO优化模型内部思考过程。我们引入PAUSE标记,在声学模糊阶段促进潜在计算,并设计感知一致性奖励以对齐推理逻辑与原始音频。跨多个基准的实验表明,HyPeR相比基线模型有显著提升,性能接近大规模模型,验证了混合感知-接地推理在鲁棒多说话人音频理解中的有效性。
原文摘要 · Abstract (English)
Recent Large Audio Language Models have demonstrated impressive capabilities in audio understanding. However, they often suffer from perceptual errors, while reliable audio reasoning is unattainable without first grounding the model's perception in structured auditory scenes. Inspired by Auditory Scene Analysis, we first introduce a Perception-Aware Question Answering (PAQA) dataset. PAQA implements a hierarchical decoupling strategy that separates speech from environmental sound and distinguishes multiple speakers, providing explicit perceptual reasoning for training. Building on this, we propose HyPeR, a two-stage Hybrid Perception-Reasoning framework. In Stage I, we finetune the model on PAQA to perceive acoustic attributes in complex audio. In Stage II, we leverage GRPO to refine the model's internal deliberation. We also introduce PAUSE tokens to facilitate latent computation during acoustically ambiguous phases and design perceptual consistency reward to align reasoning rationales with raw audio. Experiments across benchmarks demonstrate that HyPeR achieves absolute improvements over the base model, with performance comparable to large-scale models, stressing the effectiveness of hybrid perception-grounded reasoning for robust and multi-speaker audio understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。