让机器人视觉系统自己规划抓取点,提升复杂场景操作能力。
Afford-VLA: Action-Aligned Visual Planning via Internalized Affordance

- 用可学习的标记在视觉中定位交互区域,生成任务相关的可操作区域图。
- 在多个仿真和真实场景中达到顶尖性能,尤其在复杂环境任务中显著优于基线。
- 适合研究具身智能、机器人操作与视觉-语言-动作模型的开发者。
视觉-语言-动作(VLA)模型在通用机器人操作中展现出巨大潜力,但仍受限于空间推理能力不足,尤其是在复杂视觉场景中确定交互位置方面。现有方法多依赖全局几何特征、符号中间表示或外部生成的视觉信号,这些与下游动作预测耦合较弱。本文重新审视VLA系统的视觉规划,提出有效规划应具备局部性、视觉基础性、内在生成性和动作对齐性。为此,我们设计了Afford-VLA框架,将任务相关的可操作性内部化为显式的视觉规划接口。具体而言,引入可学习的<AFF>标记以查询任务相关的交互区域,从多模态特征解码可操作性掩码,并转换为紧凑嵌入,直接用于动作生成。该设计使可操作性在VLA模型内部生成并使用,形成紧密耦合的感知-动作通路。同时采用联合优化训练策略,使可操作性路径与动作预测协同改进。我们在LIBERO、LIBERO-Plus和SimplerEnv等多个仿真基准上评估,均取得一致的最先进表现,并在真实世界中验证了有效性。结果表明,将可操作性作为对齐动作的视觉规划,是一种强大的提升VLA系统的方法。
原文摘要 · Abstract (English)
Vision-language-action (VLA) models have shown strong potential for generalist robot manipulation, yet they remain limited by insufficient spatial reasoning, particularly in determining where to interact in complex visual scenes. While recent efforts introduce various forms of visual planning to address this issue, existing approaches either rely on global geometric cues, symbolic intermediate representations, or externally generated visual signals, which are often weakly coupled with downstream action prediction. In this work, we revisit visual planning in VLA systems and argue that effective planning should be local, visually grounded, internally generated, and directly aligned with action. Based on this insight, we propose Afford-VLA, a unified framework that internalizes task-conditioned affordance as an explicit visual planning interface within VLA models. Concretely, we introduce learnable <AFF> tokens to query task-relevant interaction regions, decode affordance masks from multimodal features, and convert them into compact embeddings that directly condition action generation. This design enables affordance to be both generated and utilized within the VLA, forming a tightly coupled perception-action pathway. To further support this integration, we adopt a training strategy that allows the affordance pathway to be jointly optimized with action prediction, improving its effectiveness for downstream control. We evaluate our method on multiple simulation benchmarks, including LIBERO, LIBERO-Plus, and SimplerEnv, achieving consistent state-of-the-art performance, along with strong real-world results. These findings demonstrate that internalizing affordance as action-aligned visual planning provides a powerful paradigm for improving VLA systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。