arXiv:2512.23646cs.CV2025-12被引 11

主动感知模型通过动态调用工具实现音视频细粒度理解,性能超越现有方法10%-20%。

Active Perception Agent for Omnimodal Audio-Video Understanding

  • 基于音频引导的粗到精感知,动态调度单模工具进行主动推理
  • 在三个基准上达到领先水平,无训练情况下准确率提升10%-20%
  • 适合需要高精度跨模态理解的任务,如视频内容分析与智能交互

统一音频与视觉模态的通用大语言模型虽取得进展,但在细粒度跨模态理解与多模态对齐方面仍存挑战。为此,我们提出OmniAgent,据我们所知首个完全主动感知代理,能动态协调专用单模工具以实现更精细的多模态推理。不同于以往依赖固定流程和密集帧-字幕标注的方法,本工作实现了从被动响应生成到主动多模态探询的范式转变。OmniAgent采用动态规划,按需自主调度工具调用,聚焦于任务相关的感知线索。核心是新颖的粗到精音频引导感知范式,利用音频线索定位时间事件并指导后续推理。在三个音视频理解基准上的广泛实证评估表明,OmniAgent无需训练即实现业界领先性能,准确率显著超越主流开源与闭源模型10%至20%。

原文摘要 · Abstract (English)

Omnimodal large language models have made significant strides in unifying audio and visual modalities; however, they often face challenges in fine-grained cross-modal understanding and have difficulty with multimodal alignment. To address these limitations, we introduce OmniAgent, to our best knowledge, the first fully active perception agent that dynamically orchestrates specialized unimodal tools to achieve more fine-grained omnimodal reasoning. Unlike previous works that rely on rigid, static workflows and dense frame-captioning, we demonstrate a paradigm shift from passive response generation to active multimodal inquiry. OmniAgent employs dynamic planning to autonomously orchestrate tool invocation on demand, strategically concentrating perceptual attention on task-relevant cues. Central to our approach is a novel coarse-to-fine audio-guided perception paradigm, which leverages audio cues to localize temporal events and guide subsequent reasoning. Extensive empirical evaluations on three audio-video understanding benchmarks demonstrate that OmniAgent achieves state-of-the-art performance, surpassing leading open-source and closed-source models by substantial margins of 10% - 20% accuracy without training.

多模态主动感知音视频理解推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。