arXiv:2607.00858cs.CVcs.LG2026-07

提出MoVA模型,解决长视频与文本对齐中的时间错位和语义不对称问题。

MoVA: Learning Asymmetric Dual Projections for Modular Long Video-Text Alignment

论文配图:MoVA: Learning Asymmetric Dual Projections for Modular Long Video-Text Alignment
图 1 · 摘自论文原文
  • 设计双侧非对称投影,文本侧选关键帧子空间,视频侧解耦相关视觉概念。
  • 在长视频文本任务中超越现有方法,尤其在复杂时序与长描述场景表现优异。
  • 适合研究视频-文本对齐、长序列理解或需要细粒度跨模态匹配的场景。

对比预训练推动了视频-文本对齐的发展,但模型常继承图像-文本模型(如CLIP)的关键局限,导致表征纠缠。这一问题在视频领域尤为严重:时间错位使文本仅关联特定时窗,其余帧与文本无关;语义不对称则表现为帧级视觉细节与句级概念之间稀疏、双向且非对称的相关性。该问题在短而离散的标题中引发歧义,在长而详细的描述中则加剧静态物体与动态演化的纠缠。本文建立理论条件,实现视频与文本在时间维度及不同粒度下的灵活对齐。基于此,提出MoVA(Modular Long Video-Text Alignment),学习双侧非对称投影:文本侧自适应选择关注帧的子空间,视频侧解耦文本相关的视觉概念。框架在保持全局跨模态语义的同时,能自然扩展至长文本与长视频。实证评估显示,MoVA在多项视频-文本对齐任务中优于现有方法,验证了其有效性。

原文摘要 · Abstract (English)

Contrastive pre-training has propelled video-text alignment, yet models often inherit the critical limitations of their image-text predecessors like CLIP, resulting in entangled representations. These challenges are severely exacerbated by two fundamental properties in the video domain: Temporal Misalignment, where textual descriptions often correlate only to specific, constrained temporal windows, leaving other frames text-irrelevant; and Semantic Asymmetry, which dictates a sparse, bidirectional, and non-equivalent relevance between frame-level visual details and caption-level concepts. This failure persists whether captions are short and temporally disjoint, creating ambiguity, or long and detailed, fostering entanglement between static objects and their temporal evolution. In this paper, we establish theoretical conditions that enable flexible alignment between video and text representations across the temporal dimension and at varying levels of granularity. Building on these theoretical insights, we introduce MoVA, Modular Long Video-Text Alignment, which learns dual asymmetric projections: a text-side projection that adaptively selects frame-aware subspaces of the caption, and a video-side projection that disentangles text-relevant visual concepts. Our framework ensures that the model can preserve global cross-modal semantics while disentangling evolving, frame-specific concepts and scale naturally to long captions and videos. Empirical evaluations show that MoVA outperforms existing methods in multiple video-text alignment tasks, demonstrating the effectiveness of our method.

视频-文本对齐长视频非对称投影多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。