提出新基准与智能框架,提升多模态视频推理能力
Agentic Active Omni-Modal Perception for Multi-Hop Audio-Visual Reasoning

- 构建分层记忆与动态决策循环的主动感知框架
- 在519个问题上使开源模型推理性能显著提升
- 适合需要长视频、多步推理的研究者使用
多跳音视频推理对全模态大模型仍是挑战,因相关证据稀疏、时间分散且分布于音视频流中。现有基准仅覆盖有限模态、时间片段或推理步骤。本文提出MOV-Bench基准,包含519个精心设计的问题,需跨时序多跳推理。评估显示当前全模态大模型仍难应对。为此,我们提出AOP-Agent,一种基于开源全模态大模型的高效智能体框架,通过分层多模态记忆与协同观察-反思-重规划循环,实现无需额外训练的主动感知。在MOV-Bench和OmniVideoBench上的实验表明,AOP-Agent持续提升推理表现,尤其在长视频和高复杂度问题上优势明显。
原文摘要 · Abstract (English)
Multi-hop audio-visual reasoning remains challenging for Omni-LLMs, as relevant evidence is often sparse, temporally dispersed, and distributed across both audio and visual streams. Existing benchmarks provide limited investigation of this setting, typically involving only a limited number of modalities, relevant temporal segments, or reasoning steps. In this work, we introduce MOV-Bench, a benchmark containing 519 carefully curated questions that require multi-hop reasoning over temporally dispersed audio-visual evidence. Evaluations on MOV-Bench reveal that current Omni-LLMs still struggle with multi-hop cross-modal reasoning. To address this challenge, we further propose AOP-Agent, an efficient agentic framework built on open-source Omni-LLMs for active omni-modal perception. By combining a hierarchical omni-modal memory with a collaborative observe-reflect-replan loop, AOP-Agent enables open-source Omni-LLMs to perform active perception without additional training or proprietary models. Experiments on MOV-Bench and OmniVideoBench demonstrate that AOP-Agent consistently improves reasoning performance, with particularly notable gains on long videos and reasoning-intensive questions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。