arXiv:2606.04172cs.RO2026-06被引 1

让机器人理解任务需求下的物体功能部位,实现实时操作。

Affordance2Action: Task-Conditioned Scene-level Affordance Grounding for Real-Time Manipulation

论文配图:Affordance2Action: Task-Conditioned Scene-level Affordance Grounding for Real-Time Manipulation
图 1 · 摘自论文原文
  • 基于任务条件的场景级功能部位定位,支持多区域对应。
  • 构建大规模真实场景数据集,覆盖单/多区域对应关系。
  • 适合机器人操作、具身智能等研究者使用。

任务导向的操作需要将指令与任务相关的功能部位对齐,而非仅依赖物体类别。该设定具有场景依赖性,在杂乱环境中常呈现一对多关系:同一物体在不同任务中可能具备多种交互方式,而同一任务也可能对应一个或多个有效功能区域,取决于场景布局。现有归属数据集和基准与这一设定脱节,主要聚焦抓取或物体级归属,依赖合成场景,或假设单一指令-区域对应。我们提出Affordance2Action(A2A),一个以基准为中心的学习框架,用于场景级、任务条件的功能部位对齐。核心是A2A-Bench,一个面向操作的基准,涵盖日常场景中的单区域与多区域指令对应关系,后者突显了真实多物体环境中的归属模糊性与多样性。为规模化构建,我们设计A2A-AffordGen,一种代理辅助标注流程,结合语言模型过滤、交互式部件分割、实例级掩码剔除优化、任务推理指令生成与人工验证。A2A-Bench的标注还支持多样下游应用,如实时归属对齐与归属条件操作策略。实验表明,A2A暴露出通用分割、基于视觉语言模型的对齐及归属蒸馏基线的巨大差距,同时提升任务级定位精度,并为下游操作提供有用的几何先验。所有数据集与代码将公开发布,推动开放研究。

原文摘要 · Abstract (English)

Task-conditioned manipulation requires grounding instructions to task-relevant functional parts rather than object categories. This setting is scene-dependent and often one-to-many in cluttered scenes: the same object may afford different interactions across tasks, while a single task may correspond to either one functional region or multiple valid functional regions, depending on the scene layout. Existing affordance datasets and benchmarks remain misaligned with this setting, as they typically focus on grasping or object-level affordances, rely on synthetic scenes, or assume a single instruction-region correspondence. We present Affordance2Action (A2A), a benchmark-centered learning framework for scene-level, task-conditioned part affordance grounding. At its core is A2A-Bench, a manipulation-oriented benchmark that covers both single-region and multi-region instruction correspondences in everyday scenes, with the latter highlighting the ambiguity and diversity of affordance grounding in realistic multi-object environments. To construct it at scale, we build A2A-AffordGen, an agent-assisted annotation pipeline that combines language-model filtering, interactive part segmentation, instance-level mask-out refinement, task-reasoning instruction generation, and human verification. A2A-Bench's supervision further supports diverse downstream applications, with real-time affordance grounding and affordance-conditioned manipulation policies as two representative examples. Experiments show that A2A exposes substantial gaps in generic segmentation, VLM-based grounding, and affordance distillation baselines, while improving task-level localization and providing useful spatial priors for downstream manipulation. All datasets and code will be publicly released to promote open research.

机器人操作功能归属任务条件实时控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。