无需有序关键点,用乱序3D点+图像学会打活结
RoboHitch: Learning Visual Affordance from Disordered Keypoints for Hitch Knots Tying

- 用无序3D关键点和图像学打结,不依赖拓扑顺序
- 动态图自编码器提取几何特征,融合视觉上下文
- 实测在遮挡下仍能成功打活结,适合复杂绳索操作
机器人操控柔性的线状物体(DLO)面临动态复杂和频繁自遮挡的挑战。现有打结方法通常依赖精确的拓扑状态追踪与有序关键点及显式边连接关系,易因追踪漂移和拓扑错配导致失败。为此,我们提出RoboHitch框架,仅使用人类示范中的无序3D关键点和RGB图像,学习完成活结打结任务。该方法无需显式拓扑顺序,实现更灵活的操作。我们采用动态图自编码器从未追踪的关键点中提取几何特征,同时使用卷积自编码器捕捉重要视觉上下文。双向交叉注意力机制融合多模态信息,联合预测抓取与放置的可操作性,实现对绳索状态的隐式推理,支持在遮挡条件下完成打结。真实世界实验验证了该方法的有效性与泛化能力,在存在自遮挡的情况下均成功完成活结打结。
原文摘要 · Abstract (English)
Robotic manipulation of deformable linear objects (DLOs) presents significant challenges due to complex dynamics and frequent self-occlusions. Existing robotic knot tying methods typically rely on precise topological state tracking with ordered keypoints and explicit edge connectivity. This reliance makes them prone to failures due to tracking drift and topology mismatch caused by repeated bending and crossings during knot formation.To address these limitations, we introduce RoboHitch, a novel framework that learns to perform hitch knot tying from human demonstrations using only disordered 3D keypoints and RGB images. This eliminates the need for explicit topological order, allowing for more flexible manipulation. Our method employs a dynamic Graph Autoencoder to extract geometric features from untracked keypoints, complemented by a Convolutional Autoencoder that captures essential visual context. A bidirectional cross-attention mechanism then fuses these modalities to jointly predict pick and place affordances, facilitating implicit reasoning about the rope's state and enabling knot tying under occlusion.Real-world experiments demonstrate the effectiveness and generalizability of our approach, successfully completing hitch knots in scenarios with self-occlusions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。