用视频生成模型通用生成动态物体4D场景,无需大量数据。
Choreographing a World of Dynamic Objects
- 从2D视频中提取拉格朗日运动信息,构建通用生成流程。
- 可生成多物体复杂动态,涵盖机器人操作等应用场景。
- 不依赖特定类别数据,适合广泛物理现象模拟。
真实世界中的动态物体在四维空间(3D+时间)中持续演化、变形并相互作用,形成多样化的4D场景动态。本文提出通用生成框架CHORD,用于生成动态物体与场景,并合成此类现象。传统基于规则的图形管线依赖类别特异性启发式方法,成本高且难扩展;近期学习方法通常需大规模数据集,难以覆盖所有目标类别。本方法通过借鉴视频生成模型的通用性,提出基于知识蒸馏的流程,从2D视频的欧拉表示中提取丰富的拉格朗日运动信息。该方法具有普遍性、灵活性和类别无关性。实验表明其能生成多样化的多体4D动态,优于现有方法,并成功应用于生成机器人操作策略。项目页面:https://yanzhelyu.github.io/chord
原文摘要 · Abstract (English)
Dynamic objects in our physical 4D (3D + time) world are constantly evolving, deforming, and interacting with other objects, leading to diverse 4D scene dynamics. In this paper, we present a universal generative pipeline, CHORD, for CHOReographing Dynamic objects and scenes and synthesizing this type of phenomena. Traditional rule-based graphics pipelines to create these dynamics are based on category-specific heuristics, yet are labor-intensive and not scalable. Recent learning-based methods typically demand large-scale datasets, which may not cover all object categories in interest. Our approach instead inherits the universality from the video generative models by proposing a distillation-based pipeline to extract the rich Lagrangian motion information hidden in the Eulerian representations of 2D videos. Our method is universal, versatile, and category-agnostic. We demonstrate its effectiveness by conducting experiments to generate a diverse range of multi-body 4D dynamics, show its advantage compared to existing methods, and demonstrate its applicability in generating robotics manipulation policies. Project page: https://yanzhelyu.github.io/chord
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。