arXiv:2606.02551cs.ROcs.CV2026-06

让机器人从视觉和语言理解中学会如何操作物体。

AFUN: Towards an Affordance Foundation Model for Functionality Understanding

论文配图:AFUN: Towards an Affordance Foundation Model for Functionality Understanding
图 1 · 摘自论文原文
  • 输入图像和任务描述,预测交互位置和3D动作轨迹。
  • 在8个测试集上分割准确率提升23.9%,接触点命中率提高61.3%。
  • 无需微调即可适配真实机器人,适合开放世界任务。

可操作性理解连接视觉感知与物理动作,为开放、非结构化环境中的机器人操作提供可解释接口。然而,构建一个不仅能理解交互位置与方式,还能跨环境、物体和任务泛化的可操作性基础模型,仍是长期挑战。现有方法通常仅解决部分问题:或定位相关区域但不指定动作,或预测动作但扩展性差。本文提出AFUN模型,从单张RGB-D图像和语言任务描述出发,预测任务条件下的功能掩码(何处交互)和3D接触后运动曲线(如何交互)。为支持开放世界泛化,我们构建大规模标准化数据管道,将异构的机器人、人类、仿真及真实扫描数据统一为包含语言、掩码和物体中心3D运动标签的共享可操作性模式。评估显示:在可操作性分割上,模型在4个基准的8个测试集上大幅超越所有基线,平均gIoU/cIoU提升+23.9/+26.3;在接触点预测上,命中率相比最佳基线提升12.7%–61.3%;在3D运动预测上,三个测试集均表现最优。模型无需微调即可部署于真实机器人,无需特定任务启发式,展现出对开放世界可操作性任务的强大适应能力。

原文摘要 · Abstract (English)

Affordance understanding bridges visual perception and physical action, serving as an explainable interface for robot manipulation in open and unstructured real-world environments. Yet, building an affordance foundation model that not only understands where and how the interaction should happen, but also generalizes across diverse environments, objects, and tasks, remains a long-standing research challenge. Existing methods typically address only part of this challenge, either localizing task-relevant regions without specifying executable motion, or predicting motion but with limited scalability. In this paper, we present ourmodel, a step towards an affordance foundation model for functionality understanding. From a single RGB-D observation and a language task description, ourmodel predicts a task-conditional functional mask (where to interact) and a 3D post-contact motion curve (how to interact). To support open-world generalization, we build a large-scale standardized data pipeline that converts heterogeneous robot, human, simulation, and real-world scan data into a shared affordance schema with language, masks, and object-centric 3D motion labels. We evaluate ourmodel from three aspects: for affordance segmentation, ourmodel outperforms all baselines by a large margin across 8 test sets from 4 benchmarks, improving mean gIoU/cIoU by +23.9/+26.3; for contact-point prediction, it predicts substantially more accurate points, with a 12.7--61.3% hit-rate gain over the best baseline; and for 3D motion, it achieves the best performance on all three test sets. ourmodel can be deployed for real-world robot manipulation without finetuning for robot embodiment or using task-specific heuristics, demonstrating the ability to adapt to open-world affordance tasks. Project page: https://www.zhaoningwang.com/AFUN

可操作性机器人基础模型3D动作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。