用4D重建模型精准捕捉物体部件动态,助力机器人操作
PartRM: Modeling Part-Level Dynamics with Large Cross-State Reconstruction Model
- 基于3D高斯模型构建4D动态重建框架,融合外观与几何信息
- 在超2万状态的PartDrag-4D数据集上实现当前最优部件运动预测性能
- 适用于机器人抓取与交互任务,支持快速微调且避免遗忘
随着世界模型在从当前观测和动作预测未来状态方面日益重要,准确建模部件级动态对多种应用具有重要意义。现有方法如Puppet-Master依赖微调大规模预训练视频扩散模型,因2D视频表示局限和处理速度慢,难以应用于真实场景。为此,我们提出PartRM,一种新颖的4D重建框架,可从静态物体的多视角图像中同时建模外观、几何及部件级运动。PartRM基于大规模3D高斯重建模型,利用其在静态物体上的丰富外观与几何先验知识。针对4D数据稀缺问题,我们构建了PartDrag-4D数据集,包含超过20,000个状态的部件级动态多视角观测。通过引入多尺度拖拽嵌入模块,增强模型对不同粒度动态的理解能力。为防止微调过程中的灾难性遗忘,采用分阶段训练策略,依次聚焦于运动与外观学习。实验表明,PartRM在部件级运动学习上达到新基准,并可应用于机器人操控任务。代码、数据与模型均已公开,以推动后续研究。
原文摘要 · Abstract (English)
As interest grows in world models that predict future states from current observations and actions, accurately modeling part-level dynamics has become increasingly relevant for various applications. Existing approaches, such as Puppet-Master, rely on fine-tuning large-scale pre-trained video diffusion models, which are impractical for real-world use due to the limitations of 2D video representation and slow processing times. To overcome these challenges, we present PartRM, a novel 4D reconstruction framework that simultaneously models appearance, geometry, and part-level motion from multi-view images of a static object. PartRM builds upon large 3D Gaussian reconstruction models, leveraging their extensive knowledge of appearance and geometry in static objects. To address data scarcity in 4D, we introduce the PartDrag-4D dataset, providing multi-view observations of part-level dynamics across over 20,000 states. We enhance the model's understanding of interaction conditions with a multi-scale drag embedding module that captures dynamics at varying granularities. To prevent catastrophic forgetting during fine-tuning, we implement a two-stage training process that focuses sequentially on motion and appearance learning. Experimental results show that PartRM establishes a new state-of-the-art in part-level motion learning and can be applied in manipulation tasks in robotics. Our code, data, and models are publicly available to facilitate future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。