arXiv:2508.02146cs.ROcs.CV2025-08被引 14

仅用RGB图像识别可动物体结构,实现端到端精准建模

ScrewSplat: An End-to-End Method for Articulated Object Recognition

  • 随机初始化螺旋轴,迭代优化恢复运动结构
  • 融合高斯点云实现3D重建与刚性部件分割
  • 支持零样本文本引导操作,适合机器人交互场景

可动物体识别——即识别具有可动部件的物体的几何形状与运动关节——对机器人理解日常物品(如门、笔记本)至关重要。然而,现有方法常依赖已知部件数量等强假设,或需深度图等额外输入,或包含复杂中间步骤,易引入误差,限制了实际应用。本文提出ScrewSplat,一种仅基于RGB图像的端到端方法。该方法从随机初始化螺旋轴开始,通过迭代优化恢复物体的运动学结构。结合高斯点云(Gaussian Splatting),实现3D几何重建与刚性可动部件分割同步完成。实验表明,该方法在多样化的可动物体上达到当前最优识别精度,并可进一步支持零样本、文本引导的操纵任务。项目主页:https://screwsplat.github.io。

原文摘要 · Abstract (English)

Articulated object recognition -- the task of identifying both the geometry and kinematic joints of objects with movable parts -- is essential for enabling robots to interact with everyday objects such as doors and laptops. However, existing approaches often rely on strong assumptions, such as a known number of articulated parts; require additional inputs, such as depth images; or involve complex intermediate steps that can introduce potential errors -- limiting their practicality in real-world settings. In this paper, we introduce ScrewSplat, a simple end-to-end method that operates solely on RGB observations. Our approach begins by randomly initializing screw axes, which are then iteratively optimized to recover the object's underlying kinematic structure. By integrating with Gaussian Splatting, we simultaneously reconstruct the 3D geometry and segment the object into rigid, movable parts. We demonstrate that our method achieves state-of-the-art recognition accuracy across a diverse set of articulated objects, and further enables zero-shot, text-guided manipulation using the recovered kinematic model. See the project website at: https://screwsplat.github.io.

可动物体识别端到端视觉建模机器人交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。