用合成数据提升长视频与高清视频理解能力
VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by Video Spatiotemporal Augmentation
- 通过时空拼接生成超长高清视频,构建指令跟随数据对
- 在4个长视频基准上平均提升3.3%,高清视频基准提升6.5%
- 适合需要强长视频理解能力的视觉语言模型研究者
当前大型多模态模型在处理长时长或高分辨率视频时面临显著挑战,主要源于高质量数据集的缺失。为从数据驱动角度解决此问题,我们提出VISTA——一种简单而有效的视频时空增强框架,能够从现有视频-文本数据集中合成具有长时长和高分辨率特性的视频指令跟随数据对。VISTA通过空间与时间维度组合视频,生成新合成视频,并据此生成相关问答对。基于该范式,我们设计了七种视频增强方法,构建了面向长时长与高分辨率视频理解的VISTA-400K数据集。在该数据集上微调多种视频多模态模型,在四个长视频理解挑战性基准上实现平均3.3%的性能提升。此外,我们还提出了首个全面的高分辨率视频理解基准HRVideoBench,微调模型在此基准上取得6.5%的性能增益。结果验证了该框架的有效性。
原文摘要 · Abstract (English)
Current large multimodal models (LMMs) face significant challenges in processing and comprehending long-duration or high-resolution videos, which is mainly due to the lack of high-quality datasets. To address this issue from a data-centric perspective, we propose VISTA, a simple yet effective Video Spatiotemporal Augmentation framework that synthesizes long-duration and high-resolution video instruction-following pairs from existing video-caption datasets. VISTA spatially and temporally combines videos to create new synthetic videos with extended durations and enhanced resolutions, and subsequently produces question-answer pairs pertaining to these newly synthesized videos. Based on this paradigm, we develop seven video augmentation methods and curate VISTA-400K, a video instruction-following dataset aimed at enhancing long-duration and high-resolution video understanding. Finetuning various video LMMs on our data resulted in an average improvement of 3.3% across four challenging benchmarks for long-video understanding. Furthermore, we introduce the first comprehensive high-resolution video understanding benchmark HRVideoBench, on which our finetuned models achieve a 6.5% performance gain. These results highlight the effectiveness of our framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。