从手机随手拍的视频中生成可交互的机械物体数字孪生
iTACO: Interactable Digital Twins of Articulated Objects from Casually Captured RGBD Videos
- 分步优化框架,从动态视频中推断关节参数与可动部件
- 在784段视频上验证,性能显著优于现有方法
- 适合机器人、具身智能领域快速构建实物数字模型
铰链类物体广泛存在于日常生活中。可交互的数字孪生在具身人工智能与机器人领域有广泛应用。然而,当前方法需精心采集数据,难以实现大规模、通用化获取。本文聚焦于从手持相机随意拍摄的RGBD视频中分析运动并分割部件。此类视频可通过智能手机大规模获取,但面临物体与相机同时运动、交互中严重遮挡等挑战。为此,我们提出iTACO:一种从粗到精的框架,可从动态RGBD视频中推断关节参数并分割可动部分。为评估该新设置下的方法,我们构建了一个包含784段视频、284个物体、11个类别的数据集,规模是先前工作的20倍。实验表明,iTACO在合成与真实随机拍摄的RGBD视频上均优于现有方法。
原文摘要 · Abstract (English)
Articulated objects are prevalent in daily life. Interactable digital twins of such objects have numerous applications in embodied AI and robotics. Unfortunately, current methods to digitize articulated real-world objects require carefully captured data, preventing practical, scalable, and generalizable acquisition. We focus on motion analysis and part-level segmentation of an articulated object from a casually captured RGBD video shot with a hand-held camera. A casually captured video of an interaction with an articulated object is easy to obtain at scale using smartphones. However, this setting is challenging due to simultaneous object and camera motion and significant occlusions as the person interacts with the object. To tackle these challenges, we introduce iTACO: a coarse-to-fine framework that infers joint parameters and segments movable parts of the object from a dynamic RGBD video. To evaluate our method under this new setting, we build a dataset of 784 videos containing 284 objects across 11 categories that is 20$\times$ larger than available in prior work. We then compare our approach with existing methods that also take video as input. Our experiments show that iTACO outperforms existing articulated object digital twin methods on both synthetic and real casually captured RGBD videos.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。