通过结构化推理与强化学习,提升音频视觉分割的语义理解能力
AURORA:Augmented Understanding via Structured Reasoning and Reinforcement Learning for Reference Audio-Visual Segmentation
- 采用分步推理提示机制引导模型思考
- 在保持像素精度的前提下实现推理与分割融合
- 适合需要强逻辑推理的多模态任务研究者
参考音频-视觉分割(Ref-AVS)任务要求模型结合视觉、听觉和文本线索精确定位发声对象。现有方法常缺乏真正的语义理解,倾向于记忆固定推理模式;同时,联合训练推理与分割会损害像素级精度。为此,我们提出AURORA框架,通过结构化思维链(CoT)提示机制引导模型进行分步推理,并引入新的分割特征蒸馏损失,实现推理能力的有效融合而不牺牲分割性能。为进一步提升模型的真实推理能力,设计两阶段训练策略:首先采用自修正式训练优化推理路径质量,随后通过组奖励策略优化(GRPO)的强化学习增强复杂场景下的鲁棒性。实验表明,AURORA在Ref-AVS基准上达到领先性能,并能有效泛化至未参考分割任务。
原文摘要 · Abstract (English)
Reference Audio-Visual Segmentation (Ref-AVS) tasks challenge models to precisely locate sounding objects by integrating visual, auditory, and textual cues. Existing methods often lack genuine semantic understanding, tending to memorize fixed reasoning patterns. Furthermore, jointly training for reasoning and segmentation can compromise pixel-level precision. To address these issues, we introduce AURORA, a novel framework designed to enhance genuine reasoning and language comprehension in reference audio-visual segmentation. We employ a structured Chain-of-Thought (CoT) prompting mechanism to guide the model through a step-by-step reasoning process and introduce a novel segmentation feature distillation loss to effectively integrate these reasoning abilities without sacrificing segmentation performance. To further cultivate the model's genuine reasoning capabilities, we devise a further two-stage training strategy: first, a ``corrective reflective-style training" stage utilizes self-correction to enhance the quality of reasoning paths, followed by reinforcement learning via Group Reward Policy Optimization (GRPO) to bolster robustness in challenging scenarios. Experiments demonstrate that AURORA achieves state-of-the-art performance on Ref-AVS benchmarks and generalizes effectively to unreferenced segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。