无需训练即可生成电影级长视频,精准控制镜头与风格。
CineLOG: A Training Free Approach for Cinematic Long Video Generation
- 将文本到视频生成拆解为四阶段,提升可控性
- 构建5000段高质量视频数据集,涵盖17种运镜和15类影风
- 新设计的轨迹引导过渡模块,实现多镜头自然衔接
可控视频合成是计算机视觉的核心挑战,现有模型在文本提示之外难以精确控制电影属性(如镜头运动、影片类型)。现有数据集普遍存在数据不平衡、标签噪声或仿真与真实差距大的问题。为此,我们提出CineLOG,一个包含5000段高质量、均衡且未剪辑视频片段的新数据集。每条视频均配有详细场景描述、基于标准电影分类体系的明确镜头指令及类型标签,覆盖17种多样化的镜头运动和15类电影类型。我们还设计了一套新流水线,将复杂的文本到视频生成任务分解为四个更易实现的阶段,采用成熟技术。为生成连贯的多镜头序列,提出新型轨迹引导过渡模块,实现平滑的时空插值。大量人工评估表明,该方法在遵循特定镜头与剧本指令方面显著优于当前最优端到端文本到视频模型,同时保持专业级画质。所有代码与数据已公开于https://cine-log.pages.dev。
原文摘要 · Abstract (English)
Controllable video synthesis is a central challenge in computer vision, yet current models struggle with fine grained control beyond textual prompts, particularly for cinematic attributes like camera trajectory and genre. Existing datasets often suffer from severe data imbalance, noisy labels, or a significant simulation to real gap. To address this, we introduce CineLOG, a new dataset of 5,000 high quality, balanced, and uncut video clips. Each entry is annotated with a detailed scene description, explicit camera instructions based on a standard cinematic taxonomy, and genre label, ensuring balanced coverage across 17 diverse camera movements and 15 film genres. We also present our novel pipeline designed to create this dataset, which decouples the complex text to video (T2V) generation task into four easier stages with more mature technology. To enable coherent, multi shot sequences, we introduce a novel Trajectory Guided Transition Module that generates smooth spatio-temporal interpolation. Extensive human evaluations show that our pipeline significantly outperforms SOTA end to end T2V models in adhering to specific camera and screenplay instructions, while maintaining professional visual quality. All codes and data are available at https://cine-log.pages.dev.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。