arXiv:2603.26285cs.CVcs.AI2026-03中稿 · CVPR被引 5

让视频生成更符合物理规律,通过局部动态描述提升真实感

PhysVid: Physics Aware Local Conditioning for Generative Video Models

  • 用连续帧块的物理描述作为局部条件,融合全局提示
  • 在VideoPhy上物理常识得分提升约33%,最高达8%改善
  • 适合需要高可信度视频生成的应用,如仿真与自动驾驶

生成式视频模型虽具高视觉保真度,但常违背基本物理原则,限制其在真实场景中的可靠性。先前注入物理知识的方法依赖条件输入:帧级信号具有领域特异性且时间跨度短,而全局文本提示则粗略且含噪,难以捕捉细微动态。本文提出PhysVid,一种基于局部物理感知的条件机制,针对时间连续的帧块进行标注,包含状态、相互作用与约束的物理基础描述,并通过块感知交叉注意力在训练中融合全局提示。推理时引入负向物理提示(描述局部违反物理定律的情形),引导生成避开不合理轨迹。在VideoPhy数据集上,PhysVid相较基线模型物理常识得分提升约33%,在VideoPhy2上最多提升约8%。结果表明,局部物理感知指导能显著提升生成视频的物理合理性,推动实现基于物理的视频生成模型。

原文摘要 · Abstract (English)

Generative video models achieve high visual fidelity but often violate basic physical principles, limiting reliability in real-world settings. Prior attempts to inject physics rely on conditioning: frame-level signals are domain-specific and short-horizon, while global text prompts are coarse and noisy, missing fine-grained dynamics. We present PhysVid, a physics-aware local conditioning scheme that operates over temporally contiguous chunks of frames. Each chunk is annotated with physics-grounded descriptions of states, interactions, and constraints, which are fused with the global prompt via chunk-aware cross-attention during training. At inference, we introduce negative physics prompts (descriptions of locally relevant law violations) to steer generation away from implausible trajectories. On VideoPhy, PhysVid improves physical commonsense scores by $\approx 33\%$ over baseline video generators, and by up to $\approx 8\%$ on VideoPhy2. These results show that local, physics-aware guidance substantially increases physical plausibility in generative video and marks a step toward physics-grounded video models.

视频生成物理一致性条件生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。