arXiv:2503.09994cs.CV2025-03AAAI被引 2

提升视频大模型的时序理解能力,无需额外标注数据。

TIME: Temporal-Sensitive Multi-Dimensional Instruction Tuning and Robust Benchmarking for Video-LLMs

  • 用多任务提示微调融合时序任务,不依赖昂贵时序标注。
  • 在五个维度上显著增强视频大模型的时序理解能力。
  • 新构建的基准能有效避免评估捷径,适合严肃评测。

视频大语言模型在视频问答等任务中表现优异,但其时序理解能力仍不足。为此,我们构建了一个专注提升时序理解的指令微调数据集,涵盖五个关键维度。为减少对昂贵时序标注的依赖,提出一种多任务提示微调方法,将时序敏感任务无缝融入现有指令数据集,无需额外标注。此外,开发了一种新型时序敏感视频理解基准,不仅弥补了现有基准在维度覆盖上的空白,还通过严格过滤潜在捷径,确保评估更准确。大量实验表明,该方法显著提升了视频大模型的时序理解能力,且避免了对捷径的依赖。

原文摘要 · Abstract (English)

Video large language models have achieved remarkable performance in tasks such as video question answering, however, their temporal understanding remains suboptimal. To address this limitation, we curate a dedicated instruction fine-tuning dataset that focuses on enhancing temporal comprehension across five key dimensions. In order to reduce reliance on costly temporal annotations, we introduce a multi-task prompt fine-tuning approach that seamlessly integrates temporal-sensitive tasks into existing instruction datasets without requiring additional annotations. Furthermore, we develop a novel benchmark for temporal-sensitive video understanding that not only fills the gaps in dimension coverage left by existing benchmarks but also rigorously filters out potential shortcuts, ensuring a more accurate evaluation. Extensive experimental results demonstrate that our approach significantly enhances the temporal understanding of video-LLMs while avoiding reliance on shortcuts.

视频理解时序建模指令微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。