仅用一段静态视频重建带关节的逼真物体,支持机器人仿真操作
RORA: Realistic Object Reconstruction with Articulation

- 通过用户引导的建议机制,从单个视频还原物体分段与关节结构
- 在PartNet-Mobility-v0数据集上实现高精度关节重建,支持真实物体复现
- 输出兼具视觉真实感和物理交互性的混合表示,适合机器人抓取训练
通过NeRF和3D高斯溅射(3DGS)等真实视觉表征将现实环境复刻到仿真中,已成为缓解机器人学习中模拟到现实差距的有效策略。然而,在真实到仿真的过程中实现物体关节运动仍具挑战。现有基于运动追踪或学习的关节方法在多关节复杂结构上成功率低,且需动态扫描,流程繁琐。本文提出首个端到端管道,仅需单个静态物体视频输入,结合用户引导的建议过程,重建可直接用于仿真的带关节资产。方法采用混合表示:3DGS实现逼真渲染,网格几何支持物理交互。重建流程先进行凸分解,再由用户分组实现直观部件分割,随后将3D高斯绑定至对应网格部件。自动关节建议算法基于局部边界几何计算候选关节轴,并呈现给用户以高效完成关节重建。实验表明,该方法在PartNet-Mobility-v0数据集和真实物体上均取得精确关节结果。此外,框架可用于机器人学习,已在Unreal Engine和NVIDIA Isaac Sim中部署重建资产,成功实现实时灵巧手操作任务。
原文摘要 · Abstract (English)
Replicating real-world environments into simulation by realistic visual representation like NeRF and 3D Gaussian Splatting (3DGS) has emerged as an effective strategy to reduce the sim-to-real gap in robot learning. However, implementing object articulation during the real-to-sim process is still a challenging task. Existing motion tracking or learning based articulation methods shows low success rates on complex kinematic structures having multiple joints. Furthermore, those methods require scan of dynamic motion of objects, which makes reconstruction process much complicated. In this work, we propose the first end-to-end pipeline that reconstructs simulation-ready assets with accurate articulation from a single static object video input through suggestion based human-in-the-loop process. Our approach exports a hybrid representation combining 3DGS for photorealistic rendering and mesh-based geometry for physical interaction. In the reconstruction process, our pipeline performs convex decomposition followed by user grouping for intuitive part segmentation, subsequently binding 3D Gaussians to the corresponding mesh parts. An Automatic Joint Suggestion Algorithm then calculates candidate joint axes from local boundary geometries and presents them to users for efficient articulated asset reconstruction. We have shown that our method achieves precise articulation results on partnet-mobility-v0 dataset and real objects. Additionally we presented a potential usage of our framework on robot learning, deploying the reconstructed assets in Unreal Engine and NVIDIA Isaac Sim, demonstrating real-time dexterous hand manipulation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。