用少步生成视频,让相机轨迹更自然可控。
FlashRender: Few-Step Generative Rendering via Camera-Controlled Video MeanFlow

- 通过动态对齐视频特征实现步数一致的相机控制。
- 25倍降低采样成本,视频质量媲美多步模型。
- 适合需要快速生成高质量可控视频的场景。
我们提出 FlashRender,一个可在数秒内沿目标相机轨迹重制源视频的少步生成渲染框架。发现现有多步生成渲染模型中存在依赖采样步数的相机控制误差,该误差源于离散化问题,并证明解决此不一致可显著降低去噪轨迹曲率,促进后续步数蒸馏。为此,我们引入表示变换与对齐(RETA),将源视频隐藏表示与冻结视觉几何模型的目标视频特征对齐,直接在源视频流中编码几何变换,实现采样步数一致的相机控制。随后,在 RETA 引导的低曲率去噪轨迹上,以均值流(MeanFlow)目标微调模型,有效缓解离散化误差。最后,采用基于策略的流图蒸馏,纠正固定少步采样下的自演推误差。大量实验表明,RETA、MeanFlow 和基于策略的流图蒸馏在少步生成渲染中发挥互补作用。三者结合使本方法在仅需多步基线 1/25 的采样成本下,达到相当的视频质量和几何一致性,且在分布外目标相机轨迹下仍具备更优的相机可控性。
原文摘要 · Abstract (English)
We present FlashRender, a few-step generative rendering framework that retakes a source video along a target camera trajectory in seconds. We identify sampling-step-dependent camera control as a prominent manifestation of discretization error in existing multi-step generative rendering models and show that resolving this inconsistency substantially lowers denoising trajectory curvature, facilitating subsequent step distillation. To this end, we introduce Representation Transformation and Alignment (RETA), which aligns hidden source-video representations with target-video features from a frozen visual geometry model. This directly encodes the geometric transformation within the source-video stream, enabling sampling-step-consistent camera control. We then fine-tune the model with the MeanFlow objective on the lower-curvature denoising trajectory induced by RETA, allowing the model to more effectively address discretization error. Finally, we apply on-policy flow map distillation to correct self-rollout errors under fixed few-step sampling. Extensive experiments show that RETA, MeanFlow, and on-policy flow map distillation play complementary roles in few-step generative rendering. Together, they enable our approach to match multi-step baselines in video quality and geometric consistency at 25x lower sampling cost while achieving superior camera controllability, even under out-of-distribution target camera trajectories.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。