arXiv:2503.16929cs.CVcs.AI2025-03AAAI被引 2

通过偏好学习提升视频大模型的时间推理能力

TEMPLE: Incentivizing Temporal Understanding of Video Large Language Models via Progressive Pre-SFT Alignment

  • 构建自动生成的时序偏好数据对,增强模型时间理解
  • 在多个基准上用少量数据实现性能显著提升
  • 适合希望改进视频模型时序推理的研究者

视频大语言模型(Video LLMs)通过大规模预训练和监督微调(SFT)取得显著进展,但现有方法因数据中时间对应关系弱且过度依赖下一词预测范式,导致缺乏时间监督。为此,我们提出TEMPLE(TEMporal Preference LEarning),通过直接偏好优化(DPO)系统性提升时间推理能力。为解决数据中时序信息稀缺问题,设计自动化流水线:选取时序丰富的视频,设计视频特定扰动策略,并评估模型在原始与扰动输入下的响应。此外,通过偏好学习提供额外监督信号,提出一种新型渐进式预-微调对齐策略,包含两项创新:课程学习机制逐步增加扰动难度以提高数据效率;在指令微调前应用偏好优化,激励基础时间对齐。大量实验表明,该方法在多个基准上持续提升性能,仅需少量自生成的DPO数据。结果表明,TEMPLE是SFT方法的可扩展、高效补充,为构建可靠的视频大模型开辟新路径。

原文摘要 · Abstract (English)

Video Large Language Models (Video LLMs) have achieved significant success by adopting the paradigm of large-scale pre-training followed by supervised fine-tuning (SFT). However, existing approaches struggle with temporal reasoning due to weak temporal correspondence in the data and over-reliance on the next-token prediction paradigm}, which collectively result in the absence temporal supervision. To address these limitations, we propose TEMPLE (TEMporal Preference LEarning), a systematic framework that enhances temporal reasoning capabilities through Direct Preference Optimization (DPO). To address temporal information scarcity in data, we introduce an automated pipeline for systematically constructing temporality-intensive preference pairs comprising three steps: selecting temporally rich videos, designing video-specific perturbation strategies, and evaluating model responses on clean and perturbed inputs. Complementing this data pipeline, we provide additional supervision signals via preference learning and propose a novel Progressive Pre-SFT Alignment strategy featuring two key innovations: a curriculum learning strategy which progressively increases perturbation difficulty to maximize data efficiency; and applying preference optimization before instruction tuning to incentivize fundamental temporal alignment. Extensive experiments demonstrate that our approach consistently improves Video LLM performance across multiple benchmarks with a relatively small set of self-generated DPO data. Our findings highlight TEMPLE as a scalable and efficient complement to SFT-based methods, paving the way for developing reliable Video LLMs.

视频理解偏好学习时序推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。