用合成数据训练视频多模态模型,效果逼近真实数据。
SynMulti: Synthetic-to-Real Learning for Multimodal Video Understanding

- 统一生成管道自动创建多样化的多模态视频数据。
- 合成数据训练的模型在真实数据上表现优异,超越传统方法。
- 通过问答微调提升视觉推理能力,适合研究视频理解者。
训练用于视频理解的多模态大语言模型(MLLMs)需要大规模标注数据,涵盖物体计数、问答和分割等多样化任务。然而,真实世界中收集和标注多模态视频数据成本高、速度慢,且多样性与覆盖范围有限。为此,我们提出~ extbf{SynMulti} 数据集及统一的合成数据生成流程,可自动生成无限量、丰富多样的多模态视频数据,支持多种任务格式在同一管道中实现。为进一步提升推理能力,引入基于视觉问答(VQA)的微调策略,使模型通过结构化问题回答来学习视觉内容,而非依赖描述或简单指令,从而增强视觉定位与推理。我们在三个挑战性任务上评估:视频物体计数、基于视频的视觉问答和视频物体分割。实验表明,主要在合成数据上训练的模型能有效泛化到真实数据集,常优于传统训练模型。结果凸显了统一合成数据流水线作为低成本替代真实标注的潜力,适用于多模态视频理解。
原文摘要 · Abstract (English)
Training multimodal large language models (MLLMs) for video understanding requires large-scale annotated data spanning diverse tasks such as object counting, question answering, and segmentation. However, collecting and annotating multimodal video data in real-world is costly, slow, and inherently limited in diversity and coverage. To address this challenge, we propose~\textbf{SynMulti} dataset, along with a unified synthetic data generation pipeline capable of automatically producing unlimited multimodal video data with rich and diverse supervision. Our framework supports multiple task formats within a single pipeline, enabling scalable and consistent data creation across tasks. To further enhance reasoning ability, we introduce a VQA-based fine-tuning strategy that trains models to answer structured questions about visual content rather than relying solely on captions or simple instructions. This formulation encourages deeper visual grounding and reasoning. We evaluate our approach in three challenging tasks: video object counting, video-based visual question answering, and video object segmentation. Experimental results demonstrate that models trained predominantly on synthetic data generalize effectively to real-world datasets, often outperforming traditionally trained counterparts. Our findings highlight the potential of unified synthetic data pipelines as a scalable alternative to expensive real-world annotation for multimodal video understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。