arXiv:2501.13919cs.CVcs.AI2025-01被引 34

用偏好学习提升视频模型对长视频时间定位能力

Temporal Preference Optimization for Long-Form Video Understanding

  • 通过自训练偏好数据,让模型区分准确与错误的时间定位
  • 在三个基准测试中显著提升长视频理解效果,7B模型领跑Video-MME
  • 减少人工标注依赖,适合需精准时间定位的视频分析场景

尽管视频大模型取得了显著进展,但在长视频中实现有效的时间定位仍是挑战。为此,我们提出时间偏好优化(TPO),一种基于偏好学习的后训练框架,旨在增强视频大模型的时间定位能力。TPO采用自训练方法,利用两种粒度的精选偏好数据集:局部时间定位(关注特定视频片段)和整体时间定位(捕捉全视频序列的长期依赖)。通过在这些数据集上优化,TPO显著提升了时间理解能力,同时降低了对人工标注数据的依赖。在三个长视频理解基准测试(LongVideoBench、MLVU、Video-MME)上的大量实验表明,TPO在两个领先视频大模型上均表现优异。值得注意的是,LLaVA-Video-TPO成为Video-MME上领先的7B模型,凸显TPO作为可扩展、高效的时间推理解决方案的巨大潜力。

原文摘要 · Abstract (English)

Despite significant advancements in video large multimodal models (video-LMMs), achieving effective temporal grounding in long-form videos remains a challenge for existing models. To address this limitation, we propose Temporal Preference Optimization (TPO), a novel post-training framework designed to enhance the temporal grounding capabilities of video-LMMs through preference learning. TPO adopts a self-training approach that enables models to differentiate between well-grounded and less accurate temporal responses by leveraging curated preference datasets at two granularities: localized temporal grounding, which focuses on specific video segments, and comprehensive temporal grounding, which captures extended temporal dependencies across entire video sequences. By optimizing on these preference datasets, TPO significantly enhances temporal understanding while reducing reliance on manually annotated data. Extensive experiments on three long-form video understanding benchmarks--LongVideoBench, MLVU, and Video-MME--demonstrate the effectiveness of TPO across two state-of-the-art video-LMMs. Notably, LLaVA-Video-TPO establishes itself as the leading 7B model on the Video-MME benchmark, underscoring the potential of TPO as a scalable and efficient solution for advancing temporal reasoning in long-form video understanding. Project page: https://ruili33.github.io/tpo_website.

视频理解时间定位偏好学习大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。