用视频示范控制物理行为,让生成视频更真实。
VIPER: Visual In-Context Physics Reasoning for Physically Plausible Video Generation

- 以参考视频为物理演示,提取力学规律指导生成
- 在未见数据上物理相似度提升,人类偏好更高
- 适合需要精准物理模拟的视频创作场景
现代视频生成模型虽能合成视觉吸引人且时间连贯的片段,但难以通过标准文本或图像条件控制其物理行为。核心挑战在于条件瓶颈:材料响应、接触交互、形变和运动轨迹是连续且关系复杂的物理线索,难以用语言全面描述,却可通过视频自然展示。我们提出VIPER,一种基于参考视频的视觉上下文物理推理框架,用于参考引导的图到视频生成。给定目标图像、简短提示和参考视频,VIPER将参考视频视为期望物理过程的密集视觉演示,而非外观模板。它利用多模态大语言模型(MLLM)提取参考视频中的物理线索,并通过分层训练策略引导预训练图到视频生成器,实现物理行为迁移的同时保留基础生成器的视觉先验。为此,我们构建了VIPER-19K数据集,包含材料、轨迹和物理影响标注,并筛选出参考-目标配对。在未见验证集上的实验表明,VIPER在参考视频物理相似度和人类偏好方面均优于代表性视频生成与视频作为提示基线,同时保持竞争性的一般视频质量。定性结果进一步显示,VIPER可将参考视频中的物理行为迁移到新目标场景,无需精心设计提示。
原文摘要 · Abstract (English)
Modern video generation models can synthesize visually compelling and temporally coherent clips, yet controlling their physical behavior remains difficult with standard text and image conditions. The core challenge is a conditioning bottleneck: material response, contact interaction, deformation, and motion trajectory are continuous and relational physical cues that are hard to specify exhaustively in language but can be demonstrated naturally by video. We propose VIPER, a Visual In-Context Physics Reasoning framework for reference-guided image-to-video generation. Given a target image, a brief target prompt, and a reference video, VIPER treats the reference as a dense visual demonstration of the desired physical process rather than an appearance template. It uses a Multimodal Large Language Model (MLLM) to extract reference-derived physical cues and guide a pretrained image-to-video generator through a hierarchical training strategy, enabling physical behavior transfer while preserving the visual prior of the base generator. To support this setting, we construct VIPER-19K, a curated dataset with material, trajectory, and physical-impact annotations, together with filtered reference-target pairs. Experiments on an unseen validation set show that VIPER achieves stronger reference-video physical similarity and higher human preference than representative video generation and video-as-prompt baselines, while maintaining competitive general video quality. Qualitative results further demonstrate that VIPER can transfer reference-derived physical behavior to new target scenes without requiring carefully engineered prompts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。