arXiv:2603.07136cs.RO2026-03

用关键点+扩散模型,让机器人学会通用打结

DexKnot: Generalizable Visuomotor Policy Learning for Dexterous Bag-Knotting Manipulation

  • 通过人工变形数据学习袋的通用形状表征
  • 在未见袋型上实现稳定打结,成功率超90%
  • 适合需要泛化能力的精细操作场景

打塑料袋结是日常任务,但因袋子自由度无限、物理特性复杂,对机器人极具挑战。现有方法难以泛化到未见过的袋型或形变。为此,我们提出DexKnot框架,结合关键点可表征与扩散策略,学习通用袋结策略。通过真实世界手动变形采集的关键点对应数据,学习形状无关的袋表征;对于未知袋型,通过匹配该表征识别关键点,并输入扩散变压器,基于少量人类示范生成机器人动作。该方法将观测空间降维为稀疏关键点,实现有效策略泛化。实验表明,DexKnot在多种未见袋型和形变下均实现可靠且一致的打结表现。

原文摘要 · Abstract (English)

Knotting plastic bags is a common task in daily life, yet it is challenging for robots due to the bags' infinite degrees of freedom and complex physical dynamics. Existing methods often struggle in generalization to unseen bag instances or deformations. To address this, we present DexKnot, a framework that combines keypoint affordance with diffusion policy to learn a generalizable bag-knotting policy. Our approach learns a shape-agnostic representation of bags from keypoint correspondence data collected through real-world manual deformation. For an unseen bag configuration, the keypoints can be identified by matching the representation to a reference. These keypoints are then provided to a diffusion transformer, which generates robot action based on a small number of human demonstrations. DexKnot enables effective policy generalization by reducing the dimensionality of observation space into a sparse set of keypoints. Experiments show that DexKnot achieves reliable and consistent knotting performance across a variety of previously unseen instances and deformations.

机器人操作扩散模型泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。