arXiv:2510.01433cs.ROcs.AI2025-10被引 4

用文本提示自动选关键点,让机器人更轻量高效地抓取新物体。

AFFORD2ACT: Affordance-Guided Automatic Keypoint Selection for Generalizable and Lightweight Robotic Manipulation

  • 根据文本提示从图像中自动选出最相关的2D关键点
  • 在未见过的物体上达到82%成功率,训练仅需15分钟
  • 无需感知数据或密集输入,适合快速部署到真实场景

基于视觉的机器人学习通常依赖密集图像或点云输入,计算负担重且易混入无关背景特征。现有基于关键点的方法虽能聚焦操作相关特征并保持轻量化,但要么依赖人工规则,要么任务耦合,限制了可扩展性和语义理解。为此,我们提出AFFORD2ACT,一种基于功能提示的框架,从文本提示和单张图像中提炼出最小化的语义2D关键点集。该框架采用三阶段流程:功能过滤、类别级关键点构建,以及嵌入门控机制的Transformer策略学习,以推理最相关的关键点,最终生成一个38维的状态策略,可在15分钟内完成训练,在无本体感知和密集表示的情况下实现实时运行。在多种真实世界操作任务中,该方法显著提升数据效率,在未见物体、新类别、不同背景及干扰物下均实现82%的成功率。

原文摘要 · Abstract (English)

Vision-based robot learning often relies on dense image or point-cloud inputs, which are computationally heavy and entangle irrelevant background features. Existing keypoint-based approaches can focus on manipulation-centric features and be lightweight, but either depend on manual heuristics or task-coupled selection, limiting scalability and semantic understanding. To address this, we propose AFFORD2ACT, an affordance-guided framework that distills a minimal set of semantic 2D keypoints from a text prompt and a single image. AFFORD2ACT follows a three-stage pipeline: affordance filtering, category-level keypoint construction, and transformer-based policy learning with embedded gating to reason about the most relevant keypoints, yielding a compact 38-dimensional state policy that can be trained in 15 minutes, which performs well in real-time without proprioception or dense representations. Across diverse real-world manipulation tasks, AFFORD2ACT consistently improves data efficiency, achieving an 82% success rate on unseen objects, novel categories, backgrounds, and distractors.

机器人操作关键点选择轻量化视觉学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。