让机器人理解复杂任务中物品的功能区域,实现视觉与规划的联动。
EgoAfford: Task-Oriented Affordance Grounding via Egocentric Referring Segmentation

- 通过注视视角图像和任务指令,定位下一步操作的物体、工具和目标区域。
- 构建了1.55万张标注图像的基准数据集,涵盖2000个多步场景。
- 适合研究具身智能、人机协作与多步任务规划的学者使用。
局部级功能属性定位已推动对与基础动作相关的物体区域的精准识别。将该能力扩展至复杂任务,需要将参与物体的语义角色与任务状态对齐的视觉观测及多步规划相连接。我们提出EgoAfford,一个旨在整合这三个方面的基准数据集。给定一个第一人称视角观察和高层级桌面上的任务,模型需生成后续计划,并分割出下一动作中最多三个组成部分的功能区域:直接对象、工具和目的地。EgoAfford包含约15,500张由2,000个生成的多步场景产生的经人工验证的图像,以语义对齐、任务完整的图像序列形式组织;同时提供EgoAfford-Real,共102张人工采集的图像,覆盖26个真实任务。我们还提出了EgoLens——一个具有角色特异性掩码解码器的30亿参数多模态大模型,作为该联合任务的领域内参考模型。对近期指代分割多模态大模型、商用视觉语言模型-SAM2流水线以及EgoLens的评估表明,下一步推理与动作角色条件下的部件定位存在互补性挑战。EgoLens在生成和真实观测上均展现出强大的基线性能。EgoAfford与EgoLens共同为多步桌面任务中感知与规划的联合研究提供了基础。项目主页见:https://egoafford.github.io
原文摘要 · Abstract (English)
Part-level affordance grounding has advanced the localization of functional object regions associated with elemental actions. Extending this capability to complex tasks calls for connecting the semantic roles of participating objects with task-state-aligned visual observations and multi-step planning. We introduce EgoAfford, a benchmark designed to connect these three aspects. Given an egocentric observation and a high-level tabletop task, a model must generate the remaining plan and segment the functional regions of up to three components of the next action: the direct object, instrument, and destination. EgoAfford comprises approximately 15.5k human-verified images from 2,000 generated multi-step scenes, organized as semantically aligned, task-complete image series, together with EgoAfford-Real, 102 manually captured images spanning 26 tasks. We further present EgoLens, a 3B multimodal large language model with role-specific mask decoders, as an in-domain reference model for this joint task. Evaluations of recent referring-segmentation MLLMs, commercial-VLM--SAM2 pipelines, and EgoLens highlight the complementary challenges of next-step inference and action-role-conditioned part grounding. EgoLens establishes strong reference performance on both generated and manually captured observations. Together, EgoAfford and EgoLens provide a foundation for jointly studying perception and planning in multi-step tabletop tasks. Our project page is available at: https://egoafford.github.io
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。