用任意帧生成连贯长视频,实现低成本电影级一镜到底效果
DreaMontage: Arbitrary Frame-Guided One-Shot Video Generation
- 在DiT架构中加入轻量级中间条件控制,实现任意帧精准引导
- 通过定制化数据集与强化学习优化,提升动作合理性与转场流畅度
- 采用分段自回归推理,高效生成长时序视频,适合影视创作新手
一镜到底是电影制作中一种独特而精致的美学手法,但其实际应用常受高昂成本和现实约束限制。尽管新兴视频生成模型提供了虚拟替代方案,现有方法多依赖简单的片段拼接,难以保证视觉连续性与时间一致性。本文提出DreaMontage,一个面向任意帧引导的全流程框架,可从多样用户输入生成无缝、富有表现力且时长较长的一镜到底视频。我们从三方面突破:(i) 在DiT架构中引入轻量级中间条件机制,结合自适应调优策略,有效利用基础训练数据,实现强任意帧控制能力;(ii) 构建高质量数据集并设计视觉表达微调阶段,采用定制化直接偏好优化(Tailored DPO)解决主体运动合理性与过渡平滑性问题,显著提升生成内容的成功率与可用性;(iii) 设计分段自回归(SAR)推理策略,在保持内存效率的同时支持长序列生成。大量实验表明,该方法在视觉冲击力与时间连贯性上均表现优异,兼具计算高效性,使用户能将零散视觉素材转化为生动、统一的一镜到底电影体验。
原文摘要 · Abstract (English)
The "one-shot" technique represents a distinct and sophisticated aesthetic in filmmaking. However, its practical realization is often hindered by prohibitive costs and complex real-world constraints. Although emerging video generation models offer a virtual alternative, existing approaches typically rely on naive clip concatenation, which frequently fails to maintain visual smoothness and temporal coherence. In this paper, we introduce DreaMontage, a comprehensive framework designed for arbitrary frame-guided generation, capable of synthesizing seamless, expressive, and long-duration one-shot videos from diverse user-provided inputs. To achieve this, we address the challenge through three primary dimensions. (i) We integrate a lightweight intermediate-conditioning mechanism into the DiT architecture. By employing an Adaptive Tuning strategy that effectively leverages base training data, we unlock robust arbitrary-frame control capabilities. (ii) To enhance visual fidelity and cinematic expressiveness, we curate a high-quality dataset and implement a Visual Expression SFT stage. In addressing critical issues such as subject motion rationality and transition smoothness, we apply a Tailored DPO scheme, which significantly improves the success rate and usability of the generated content. (iii) To facilitate the production of extended sequences, we design a Segment-wise Auto-Regressive (SAR) inference strategy that operates in a memory-efficient manner. Extensive experiments demonstrate that our approach achieves visually striking and seamlessly coherent one-shot effects while maintaining computational efficiency, empowering users to transform fragmented visual materials into vivid, cohesive one-shot cinematic experiences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。