无需训练即可精准控制视频生成相机视角,提升视觉质量与轨迹一致性。
$h$-control: Training-Free Camera Control via Block-Conditional Gibbs Refinement

- 通过块条件伪吉布斯精修,在不更新模型的前提下优化未观测潜空间。
- 在RealEstate10K和DAVIS上优于所有七种训练自由与训练依赖方法,FVD最低。
- 适用于需要高精度相机控制的视频生成场景,尤其适合无训练资源用户。
针对预训练流匹配视频生成器的无训练相机控制问题,本质是部分观测下的逆问题:深度扭曲引导视频仅提供部分潜空间的噪声证据,采样器需结合预训练先验进行推断。现有方法难以平衡轨迹遵循与视觉质量,且启发式引导强度调节缺乏鲁棒性。本文提出 $h$-control,通过重构采样器结构实现突破:每个外层硬替换引导步骤均增加一个内层块条件伪吉布斯精修,对同一噪声水平下未观测补集进行优化,并具备收敛至部分观测条件数据分布的理论保证。为加速高维视频潜空间的收敛,利用其条件局部性,将未观测部分划分为3D块,每块由自适应冻结机制跟踪的定制混合指示器维护。在RealEstate10K与DAVIS数据集上,$h$-control在所有七个训练自由与训练依赖对比方法中取得最优FVD表现,且在所有报告指标上超越每种训练自由基线。
原文摘要 · Abstract (English)
Training-free camera control for pretrained flow-matching video generators is a partial-observation inverse problem: a depth-warped guidance video supplies noisy evidence on a subset of latent sites, which the sampler must reconcile with the pretrained prior. Existing methods struggle to balance the trade-off between trajectory adherence and visual quality and the heuristic guidance-strength tuning lacks robustness. We propose \textbf{$h$-control}, which resolves this dilemma through a structural change to the sampler: each outer hard-replacement guidance step is augmented with an inner-loop \emph{block-conditional pseudo-Gibbs refinement} on the unobserved complement at the same noise level, with provable convergence to the partial-observation conditional data law. To accelerate convergence on high-dimensional video latents, we exploit their conditional locality, partitioning the unobserved complement into 3D patches, each tracked by a custom mixing indicator that adaptively freezes converged patches. On RealEstate10K and DAVIS, \textbf{$h$-control} attains the best FVD against all seven training-free and training-based competitors, outperforming every training-free baseline on every reported metric.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。