通过双策略提升音视频联合理解能力,模型更懂多模态线索。
OmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention
- 基于自监督学习的查询聚焦定位,强化跨模态对齐。
- 对比学习驱动的模态注意力融合,显著提升推理效果。
- 适合需要强多模态理解的视频分析任务研究者。
人类通过多种感官协同感知世界以获得整体理解,但现有全模态视频模型在音视频理解任务中仍面临重大挑战。本文提出 OmniVideo-R1,一种增强型多模态推理框架,通过两种关键策略实现跨模态深度理解:(1) 基于自监督学习的查询密集型定位;(2) 基于对比学习的模态注意力融合。在多个基准测试上的大量实验表明,OmniVideo-R1 持续优于强大基线模型,验证了其有效性与出色的泛化能力。
原文摘要 · Abstract (English)
While humans perceive the world through diverse modalities that operate synergistically to support a holistic understanding of their surroundings, existing omnivideo models still face substantial challenges on audio-visual understanding tasks. In this paper, we propose OmniVideo-R1, a novel reinforced framework that improves mixed-modality reasoning. OmniVideo-R1 empowers models to "think with omnimodal cues" by two key strategies: (1) query-intensive grounding based on self-supervised learning paradigms; and (2) modality-attentive fusion built upon contrastive learning paradigms. Extensive experiments on multiple benchmarks demonstrate that OmniVideo-R1 consistently outperforms strong baselines, highlighting its effectiveness and robust generalization capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。