arXiv:2512.16250cs.AIcs.MA2025-12被引 4

提出音频视觉多说话人理解新基准与对齐框架,提升模型对话推理能力。

AMUSE: Audio-Visual Benchmark and Alignment Framework for Agentic Multi-Speaker Understanding

  • 构建需规划、定位、反思的多说话人任务,评估模型在三种模式下的表现。
  • 当前模型在多说话人场景下准确率低,行为不一致,尤其在主动推理时更差。
  • 提出高效对齐框架RAFT,结合奖励优化与自评估,提升39.52%准确率。

近期多模态大语言模型(如GPT-4o和Qwen3-Omni)虽具备强大感知能力,但在需要追踪发言者、维持角色并跨时间锚定事件的多说话人对话场景中表现不佳。这类场景是多模态音视频理解的核心,广泛应用于对话视频助手和会议分析等任务。我们提出AMUSE基准,围绕具有内在代理特性的任务设计,要求模型将复杂音视频交互分解为规划、定位与反思步骤。该基准评估多模态大模型在零样本、引导式和代理式三种模式下的表现,涵盖六类任务,包括时空说话人定位和多模态对话摘要。结果显示,当前模型在所有模式下均表现出弱多说话人推理能力,且非代理与代理评估下行为不一致。受任务天然代理属性及大语言模型代理进展启发,我们提出RAFT框架:通过奖励优化结合内在多模态自我评估作为奖励信号,并采用选择性参数适配实现数据与参数高效更新。使用RAFT后,在该基准上最高实现39.52%相对准确率提升。AMUSE与RAFT共同构建了评估与增强多模态模型代理推理能力的实用平台。

原文摘要 · Abstract (English)

Recent multimodal large language models (MLLMs) such as GPT-4o and Qwen3-Omni show strong perception but struggle in multi-speaker, dialogue-centric settings that demand agentic reasoning tracking who speaks, maintaining roles, and grounding events across time. These scenarios are central to multimodal audio-video understanding, where models must jointly reason over audio and visual streams in applications such as conversational video assistants and meeting analytics. We introduce AMUSE, a benchmark designed around tasks that are inherently agentic, requiring models to decompose complex audio-visual interactions into planning, grounding, and reflection steps. It evaluates MLLMs across three modes zero-shot, guided, and agentic and six task families, including spatio-temporal speaker grounding and multimodal dialogue summarization. Across all modes, current models exhibit weak multi-speaker reasoning and inconsistent behavior under both non-agentic and agentic evaluation. Motivated by the inherently agentic nature of these tasks and recent advances in LLM agents, we propose RAFT, a data-efficient agentic alignment framework that integrates reward optimization with intrinsic multimodal self-evaluation as reward and selective parameter adaptation for data and parameter efficient updates. Using RAFT, we achieve up to 39.52\% relative improvement in accuracy on our benchmark. Together, AMUSE and RAFT provide a practical platform for examining agentic reasoning in multimodal models and improving their capabilities.

多模态语音理解代理推理基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。