arXiv:2603.13770cs.CV2026-03被引 2

让视频生成符合物理规律,解决动态画面不连贯问题。

PhysAlign: Physics-Coherent Image-to-Video Generation through Feature and 3D Representation Alignment

  • 通过刚体模拟构建带精确物理标注的合成数据集
  • 融合3D几何约束与时空关系对齐,实现物理一致性
  • 适合需要真实物理推理的机器人和影视生成场景

视频扩散模型(VDMs)在模拟动态场景方面潜力巨大,但现有模型常生成违背基本物理直觉的时序不一致内容,严重限制其应用。本文提出PhysAlign框架,通过基于刚体模拟的全可控合成数据生成管道,构建了高精度、细粒度物理与3D标注的数据集。利用该数据,PhysAlign通过耦合显式3D几何约束与基于Gram矩阵的时空关系对齐机制,建立统一的物理潜在空间,从视频基础模型中提取运动先验。大量实验表明,PhysAlign在需复杂物理推理与时间稳定性的任务中显著优于现有VDMs,且不牺牲零样本视觉质量。该方法为视觉合成与刚体运动学之间建立了桥梁,推动真正物理驱动的视频生成落地。项目主页:https://physalign.github.io/PhysAlign。

原文摘要 · Abstract (English)

Video Diffusion Models (VDMs) offer a promising approach for simulating dynamic scenes and environments, with broad applications in robotics and media generation. However, existing models often generate temporally incoherent content that violates basic physical intuition, significantly limiting their practical applicability. We propose PhysAlign, an efficient framework for physics-coherent image-to-video (I2V) generation that explicitly addresses this limitation. To overcome the critical scarcity of physics-annotated videos, we first construct a fully controllable synthetic data generation pipeline based on rigid-body simulation, yielding a highly-curated dataset with accurate, fine-grained physics and 3D annotations. Leveraging this data, PhysAlign constructs a unified physical latent space by coupling explicit 3D geometry constraints with a Gram-based spatio-temporal relational alignment that extracts kinematic priors from video foundation models. Extensive experiments demonstrate that PhysAlign significantly outperforms existing VDMs on tasks requiring complex physical reasoning and temporal stability, without compromising zero-shot visual quality. PhysAlign shows the potential to bridge the gap between raw visual synthesis and rigid-body kinematics, establishing a practical paradigm for genuinely physics-grounded video generation. The project page is available at https://physalign.github.io/PhysAlign.

视频生成物理对齐扩散模型3D表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。