将任意输入转化为结构化描述,提升视频生成的可控性与质量
Any2Caption:Interpreting Any Condition to Caption for Controllable Video Generation
- 用多模态大模型解析文本、图像、视频等多元输入为结构化描述
- 在33.7万条数据上训练,显著提升现有视频模型的可控性与画质
- 适合需要精准控制视频生成内容的研究者与开发者
为解决当前视频生成领域中用户意图理解不准的瓶颈,我们提出Any2Caption框架,实现任意条件下的可控视频生成。核心思想是将条件解析与视频合成过程解耦。通过现代多模态大语言模型(MLLMs),Any2Caption可将文本、图像、视频及区域、运动、相机位姿等专用提示,转化为密集且结构化的描述,为视频生成器提供更优引导。我们还构建了Any2CapIns数据集,包含33.7万实例与407万条件,用于任意条件到描述的指令微调。全面评估表明,该系统在多种现有视频生成模型中均显著提升了可控性与视频质量。
原文摘要 · Abstract (English)
To address the bottleneck of accurate user intent interpretation within the current video generation community, we present Any2Caption, a novel framework for controllable video generation under any condition. The key idea is to decouple various condition interpretation steps from the video synthesis step. By leveraging modern multimodal large language models (MLLMs), Any2Caption interprets diverse inputs--text, images, videos, and specialized cues such as region, motion, and camera poses--into dense, structured captions that offer backbone video generators with better guidance. We also introduce Any2CapIns, a large-scale dataset with 337K instances and 407K conditions for any-condition-to-caption instruction tuning. Comprehensive evaluations demonstrate significant improvements of our system in controllability and video quality across various aspects of existing video generation models. Project Page: https://sqwu.top/Any2Cap/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。