arXiv:2503.01092cs.CVcs.RO2025-03中稿 · IROS 2025被引 3

仅用一次样本,让机器人识别变形物体的可操作区域。

One-Shot Affordance Grounding of Deformable Objects in Egocentric Organizing Scenes

  • 通过结构增强模块提升对物体内部结构的理解能力。
  • 在真实场景数据集上,各项指标提升超2.9%以上。
  • 适合需要快速学习新变形物体的机器人应用。

机器人在操作变形物体时面临组件属性不确定、配置多样、视觉干扰和提示模糊等挑战,导致感知与控制困难。为此,本文提出一种基于单次样本的变形物体可操作性定位方法(OS-AGDO),使机器人能仅凭少量样本即可识别颜色和形状各异的未见变形物体。首先引入变形物体语义增强模块(DefoSEM),强化对物体内部结构的分层理解,提升局部特征识别精度,即使在组件信息弱的情况下仍有效。其次提出基于ORB的关键点融合模块(OEKFM),利用几何约束优化关键部件特征提取,增强对多样性与视觉干扰的适应性。此外设计基于图像与任务上下文的实例条件提示,缓解因提示词导致的区域歧义。为验证方法,构建了包含15类常见变形物体及其对应整理动作的真实世界数据集AGDDO15。实验表明,本方法在KLD、SIM、NSS三项指标上分别提升6.2%、3.2%、2.9%,且具有强泛化能力。源代码与数据集已公开于https://github.com/Dikay1/OS-AGDO。

原文摘要 · Abstract (English)

Deformable object manipulation in robotics presents significant challenges due to uncertainties in component properties, diverse configurations, visual interference, and ambiguous prompts. These factors complicate both perception and control tasks. To address these challenges, we propose a novel method for One-Shot Affordance Grounding of Deformable Objects (OS-AGDO) in egocentric organizing scenes, enabling robots to recognize previously unseen deformable objects with varying colors and shapes using minimal samples. Specifically, we first introduce the Deformable Object Semantic Enhancement Module (DefoSEM), which enhances hierarchical understanding of the internal structure and improves the ability to accurately identify local features, even under conditions of weak component information. Next, we propose the ORB-Enhanced Keypoint Fusion Module (OEKFM), which optimizes feature extraction of key components by leveraging geometric constraints and improves adaptability to diversity and visual interference. Additionally, we propose an instance-conditional prompt based on image data and task context, which effectively mitigates the issue of region ambiguity caused by prompt words. To validate these methods, we construct a diverse real-world dataset, AGDDO15, which includes 15 common types of deformable objects and their associated organizational actions. Experimental results demonstrate that our approach significantly outperforms state-of-the-art methods, achieving improvements of 6.2%, 3.2%, and 2.9% in KLD, SIM, and NSS metrics, respectively, while exhibiting high generalization performance. Source code and benchmark dataset are made publicly available at https://github.com/Dikay1/OS-AGDO.

机器人操作变形物体单次学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。