让视频生成模型自己优化生成过程,提升物理真实感。
Self-Refining Video Sampling
- 用预训练模型自身作为去噪器,在推理时迭代优化视频帧。
- 在多个视频生成模型上实现70%以上人类偏好度提升。
- 通过自一致性判断选择性优化区域,避免过度修正产生伪影。
当前视频生成模型在复杂物理动态方面仍存在不足,难以达到物理真实性。现有方法依赖外部验证器或额外数据训练,计算成本高且对细微运动捕捉能力有限。本文提出自精炼视频采样,利用大规模预训练视频生成模型作为自身自精炼器。将生成器视为去噪自编码器,在推理阶段实现无外部验证器、无需额外训练的迭代内循环优化。进一步提出基于不确定性的精炼策略,依据自一致性选择性优化区域,防止因过度精炼引发的伪影。在多个先进视频生成模型上的实验表明,该方法显著提升了运动连贯性与物理一致性,人类偏好度超过70%,优于默认采样器和基于引导的采样器。
原文摘要 · Abstract (English)
Modern video generators still struggle with complex physical dynamics, often falling short of physical realism. Existing approaches address this using external verifiers or additional training on augmented data, which is computationally expensive and still limited in capturing fine-grained motion. In this work, we present self-refining video sampling, a simple method that uses a pre-trained video generator trained on large-scale datasets as its own self-refiner. By interpreting the generator as a denoising autoencoder, we enable iterative inner-loop refinement at inference time without any external verifier or additional training. We further introduce an uncertainty-aware refinement strategy that selectively refines regions based on self-consistency, which prevents artifacts caused by over-refinement. Experiments on state-of-the-art video generators demonstrate significant improvements in motion coherence and physics alignment, achieving over 70% human preference compared to the default sampler and guidance-based sampler.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。