用语言控制深度感知的分层物理动画,让静态图动得更真实。
PhysLayer: Language-Guided Layered Animation with Depth-Aware Physics

- 基于语言理解将图像分解为深度分层,指导物理模拟。
- 在2D基础上实现深度运动与透视缩放,提升交互真实性。
- 适合需要可控、真实动态效果的视频生成场景。
现有图像到视频生成方法常产生不符合物理规律的动作,且难以精确控制物体动态。尽管先前工作引入了物理模拟器,但仍局限于二维平面运动,无法捕捉深度感知的空间交互。我们提出PhysLayer,一种新型框架,实现语言引导的深度感知分层动画。该框架包含三个关键组件:首先,利用视觉基础模型的语言引导场景理解模块,根据物体组成、材质属性和物理参数将场景分解为基于深度的层次;其次,深度感知的分层物理模拟扩展了二维刚体动力学,加入深度运动和透视一致的缩放,无需完整3D重建即可实现更真实的物体交互;第三,物理引导的视频合成模块将模拟轨迹与场景感知光照融合,实现时间连贯的结果。实验表明,相比基线,在CLIP相似度(+2.2%)、FID得分(+9.3%)和运动FID(+3%)上均有提升,人工评估显示物理合理性提升24%,文本-视频对齐度提升35%。本方法在物理真实性和计算效率间取得实用平衡,适用于可控图像动画。
原文摘要 · Abstract (English)
Existing image-to-video generation methods often produce physically implausible motions and lack precise control over object dynamics. While prior approaches have incorporated physics simulators, they remain confined to 2D planar motions and fail to capture depth-aware spatial interactions. We introduce PhysLayer, a novel framework enabling language-guided, depth-aware layered animation of static images. PhysLayer consists of three key components: First, a language-guided scene understanding module that utilizes vision foundation models to decompose scenes into depth-based layers by analyzing object composition, material properties, and physical parameters. Second, a depth-aware layered physics simulation that extends 2D rigid-body dynamics with depth motion and perspective-consistent scaling, enabling more realistic object interactions without requiring full 3D reconstruction. Third, a physics-guided video synthesis module that integrates simulated trajectories with scene-aware relighting for temporally coherent results. Experimental results demonstrate improvements in CLIP-Similarity (+2.2\%), FID score (+9.3\%), and Motion-FID (+3\%), with human evaluation showing enhanced physical plausibility (+24\%) and text-video alignment (+35\%). Our approach provides a practical balance between physical realism and computational efficiency for controllable image animation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。