用合成数据训练视频模型,性能媲美真实数据。
LLaVA-Video: Video Instruction Tuning With Synthetic Data
- 构建17.8万条合成视频指令数据,涵盖描述与问答任务。
- 在多个视频基准上表现优异,超越现有方法。
- 适合研究视频多模态、指令微调的开发者使用。
视频大模型的发展受限于高质量网络原始数据的获取。为此,我们提出一种新方法:构建高质量合成数据集用于视频指令跟随,即LLaVA-Video-178K。该数据集包含详细描述、开放式问答(QA)和多项选择题问答等关键任务。结合现有视觉指令微调数据训练,我们推出了新视频大模型LLaVA-Video。实验表明,该模型在多个视频基准上表现强劲,验证了数据集的有效性。我们将发布数据集、生成流程及模型权重。
原文摘要 · Abstract (English)
The development of video large multimodal models (LMMs) has been hindered by the difficulty of curating large amounts of high-quality raw data from the web. To address this, we propose an alternative approach by creating a high-quality synthetic dataset specifically for video instruction-following, namely LLaVA-Video-178K. This dataset includes key tasks such as detailed captioning, open-ended question-answering (QA), and multiple-choice QA. By training on this dataset, in combination with existing visual instruction tuning data, we introduce LLaVA-Video, a new video LMM. Our experiments demonstrate that LLaVA-Video achieves strong performance across various video benchmarks, highlighting the effectiveness of our dataset. We plan to release the dataset, its generation pipeline, and the model checkpoints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。