通过时间重排任务提升多模态模型的视频音频理解与协同推理能力
OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering

- 设计时间重排代理任务,分三策略协同视听信号进行跨模态融合
- 在15个基准上实现视频、音频及协同推理性能显著提升
- 适合需要自监督多模态学习的科研与工业场景
为将强化学习后训练范式拓展至多模态模型,以同步增强视频-音频理解与协同推理能力,我们提出OmniJigsaw——一种基于时间重排代理任务的通用自监督框架。该框架围绕打乱的音视频片段的时序重建,通过三种策略协同视觉与听觉信号:联合模态融合、样本级模态选择与片段级模态掩码。针对代理任务有效性取决于拼图质量的问题,设计两级粗到细的数据过滤流程,实现对海量无标注多模态数据的高效适配。分析发现联合模态融合存在“双模态捷径现象”,而细粒度片段级模态掩码可缓解此问题并优于样本级模态选择。在15个基准上的广泛评估显示,视频、音频及协同推理性能均有显著提升,验证了OmniJigsaw作为可扩展的自监督多模态学习范式的有效性。
原文摘要 · Abstract (English)
To extend the reinforcement learning post-training paradigm to omni-modal models for concurrently bolstering video-audio understanding and collaborative reasoning, we propose OmniJigsaw, a generic self-supervised framework built upon a temporal reordering proxy task. Centered on the chronological reconstruction of shuffled audio-visual clips, this paradigm strategically orchestrates visual and auditory signals to compel cross-modal integration through three distinct strategies: Joint Modality Integration, Sample-level Modality Selection, and Clip-level Modality Masking. Recognizing that the efficacy of such proxy tasks is fundamentally tied to puzzle quality, we design a two-stage coarse-to-fine data filtering pipeline, which facilitates the efficient adaptation of OmniJigsaw to massive unannotated omni-modal data. Our analysis reveals a ``bi-modal shortcut phenomenon'' in joint modality integration and demonstrates that fine-grained clip-level modality masking mitigates this issue while outperforming sample-level modality selection. Extensive evaluations on 15 benchmarks show substantial gains in video, audio, and collaborative reasoning, validating OmniJigsaw as a scalable paradigm for self-supervised omni-modal learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。