arXiv:2606.06155cs.ROcs.CV2026-06被引 2

让机器人更懂物体能做什么,从而精准执行指令。

AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding

论文配图:AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding
图 1 · 摘自论文原文
  • 用物体可操作性作为中间表示,连接视觉、语言和动作
  • 在仿真与真实场景中均实现高效精准的操控表现
  • 适合需要理解物体功能的智能机器人研发者

视觉-语言-动作(VLA)模型利用预训练视觉-语言模型(VLM)中的丰富世界知识,实现基于指令的机器人操作。然而,VLM语义空间与具身控制策略之间的结构不匹配常导致感知-动作映射不精确。为此,我们提出AffordanceVLA,一种统一框架,通过引入结构化的可操作性预测作为任务导向的中间表示,建立更精确、鲁棒的感知-动作映射。具体地,通过三个互补模块逐步建模操作先验:1) Which2Act,基于视觉隐变量预测实现以物体为中心的定位,抑制干扰;2) Where2Act,通过可操作性图估计实现2D交互定位;3) How2Act,进行3D几何推理以指导操作策略。这些可操作性线索提供空间上对齐、语义条件化且与动作耦合的中间表示,自然桥接视觉、语言与动作。我们将这些模块集成到混合式Transformer(MoT)架构中,采用三阶段训练策略与渐进式数据课程进行训练。为克服机器人数据集中密集可操作性标签稀缺的问题,我们还开发了鲁棒的自动化数据增强流水线。大量仿真与真实世界实验表明,AffordanceVLA在多样化操作场景中均表现出色。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models leverage the rich world knowledge of pretrained vision-language models (VLMs) to enable instruction-following robotic manipulation. However, the structural mismatch between VLM semantic spaces and embodied control policies often hinders the learning of precise perception--action mappings. To address this challenge, we propose \textbf{AffordanceVLA}, a unified framework that introduces structured affordance forecasting as a task-oriented intermediate representation to establish a more precise and robust perception--action mapping. Specifically, we progressively model manipulation priors through three complementary components: 1) \textbf{Which2Act} for object-centric grounding via visual latent prediction to suppress distractions; 2) \textbf{Where2Act} for 2D interaction localization via affordance map estimation; and 3) \textbf{How2Act} for 3D geometric reasoning to guide manipulation policies. These affordance cues provide spatially grounded, semantically conditioned, and action-coupled intermediate representations, thereby naturally bridging vision, language and action. We integrate these modules into a Mixture-of-Transformer (MoT) architecture with specialized experts and train the model using a three-stage training strategy with a progressive data curriculum. To overcome the scarcity of dense affordance labels in robotic datasets, we also develop a robust automated data augmentation pipeline. Extensive experiments on simulation and real-world demonstrate that AffordanceVLA achieves strong performance across diverse manipulation scenarios.

机器人操作视觉语言动作可操作性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。