arXiv:2506.12009cs.CV2025-06被引 6

用自动化生成75万张3D交互热图,让机器人学会在任意物体上找可操作位置。

Affogato: Open-Vocabulary Affordance Grounding with Automated Data Generation at Scale

  • 用基础模型自动构建大规模3D交互热图数据集,无需人工标注。
  • 在75万张热图上预训练,对未见物体和操作类型提升显著。
  • 适合做开放词汇交互理解的科研与机器人研发人员。

交互定位是具身智能体的核心能力,但受限于数据:人工标注成本高,现有数据集仅覆盖有限类别。本文提出Affogato框架,构建了包含75万张3D交互热图与自然语言查询的大型数据集Affogato-750K,通过全自动化流程生成,无须人工标注,涵盖远超以往的数据类别多样性。为确保评估可靠性,额外提供5000对人工验证测试样本。同时提出Espresso-3D与Espresso-2D两个统一架构的轻量级模型。在Affogato-750K上预训练后,两种模型及已有方法均获得性能提升,尤其在未见物体与操作类别上效果最佳,证明该数据集能提供跨架构的广泛可迁移监督信号。

原文摘要 · Abstract (English)

Affordance grounding aims to localize where to interact with an object, a fundamental capability for embodied agents. Yet progress is bottlenecked by data: manual annotation is prohibitively expensive and confines existing datasets to a narrow set of predefined object and affordance categories. We introduce Affogato, a framework for open-vocabulary affordance grounding centered on Affogato-750K, a large-scale dataset of 750K 3D affordance heatmaps paired with natural language queries. We build it with a fully automated pipeline that orchestrates foundation models to generate them at scale without human labeling. It covers significantly more diverse categories than any existing dataset. For reliable evaluation, we further provide 5K human-verified test pairs. We also present Espresso-3D and Espresso-2D, simple yet effective models with a unified architecture across both modalities. Pretraining on Affogato-750K improves both Espresso and prior methods and yields the largest gains on unseen object and affordance categories, showing that it provides broadly transferable supervision across architectures.

交互定位3D感知自动化标注开放词汇

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。