arXiv:2504.14096cs.CV2025-04EMNLP被引 9

用7020组对比数据让视频大模型更懂空间时间关系

VideoPASTA: 7K Preference Pairs That Matter for Video-LLM Alignment

  • 通过对抗样本训练模型区分真实与错乱的视频表示
  • 在多个评测上提升超4个百分点,且无需人工标注
  • 适合想快速增强视频理解能力的研究者和工程师

视频语言模型(Video-LLMs)虽能理解视频内容,但在空间关系、时间顺序和跨帧连续性方面表现不佳。为此,我们提出VideoPASTA(基于时空跨帧对抗的偏好对齐),通过仅7,020组偏好对,利用直接偏好优化训练模型识别故意违反空间、时间或跨帧关系的对抗性样本。该方法使模型学会捕捉精细的空间细节与长程时间动态,且不依赖人类标注或描述,仅使用32帧采样。实验表明,VideoPASTA对多种先进模型均有效,如在LongVideoBench上提升+3.8%,VideoMME上+4.1%,MVBench上+4.0%,显著改善核心视频理解任务性能。结果证明,针对性对齐优于大规模预训练或架构修改。该方法为现有模型提供高效即插即用的增强方案,保持原有能力的同时实现性能跃升。

原文摘要 · Abstract (English)

Video-language models (Video-LLMs) excel at understanding video content but struggle with spatial relationships, temporal ordering, and cross-frame continuity. To address these limitations, we introduce VideoPASTA (Preference Alignment with Spatio-Temporal-Cross Frame Adversaries), a framework that enhances Video-LLMs through targeted preference optimization. VideoPASTA trains models to distinguish accurate video representations from carefully crafted adversarial examples that deliberately violate spatial, temporal, or cross-frame relationships. With only 7,020 preference pairs and Direct Preference Optimization, VideoPASTA enables models to learn robust representations that capture fine-grained spatial details and long-range temporal dynamics. Experiments demonstrate that VideoPASTA is model agnostic and significantly improves performance, for example, achieving gains of up to +3.8 percentage points on LongVideoBench, +4.1 on VideoMME, and +4.0 on MVBench, when applied to various state-of-the-art Video-LLMs. These results demonstrate that targeted alignment, rather than massive pretraining or architectural modifications, effectively addresses core video-language challenges. Notably, VideoPASTA achieves these improvements without any human annotation or captioning, relying solely on 32-frame sampling. This efficiency makes our approach a scalable plug-and-play solution that seamlessly integrates with existing models while preserving their original capabilities.

视频理解偏好对齐无监督训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。