arXiv:2605.17517cs.RO2026-05被引 3

让机器人模型学会关注物体可操作区域,提升复杂环境下的抓取成功率。

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment

论文配图:AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment
图 1 · 摘自论文原文
  • 通过隐式对齐机制,将操作功能信息融入视觉表征。
  • 在仿真和真实场景中均达到顶尖性能,成功率达92.3%。
  • 无需额外标注或模块,兼顾准确率与推理效率,适合实际部署。

视觉-语言-动作(VLA)模型在通用机器人操作中展现出巨大潜力,但其视觉表征常被全局物体外观主导,难以聚焦任务相关的功能交互区域,限制了在非结构化环境中的鲁棒性。现有基于功能的方法通常依赖显式掩码注入或外部感知模块,需额外标注并引入级联感知误差和推理开销。为此,我们提出AffordVLA,一种通过隐式表示对齐将操作中心的功能感知内化到VLA视觉表征中的增强框架。具体而言,构建零样本功能教师,从RGB观测和语言指令中提取任务条件化的功能视觉表征;再将VLA的中间视觉表征与教师提取的功能表征对齐,从而隐式注入操作中心的功能感知,提升动作准确性。大量仿真与真实世界实验表明,AffordVLA及其功能教师达到当前最优表现,优于多个强基线。消融分析显示,AffordVLA有效重塑VLA视觉表征,同时保持高效推理,显著提升操作成功率与训练效率。

原文摘要 · Abstract (English)

Recent advances in Vision-Language-Action (VLA) models have shown strong potential for general-purpose robotic manipulation. However, the visual representations of most VLA models are often dominated by global object appearance and struggle to focus on task-relevant functional interaction regions, which limits their robustness in unstructured environments. Existing affordance-based methods typically rely on explicit mask injection or external perception modules, requiring additional annotations while introducing cascading perception errors and inference overhead. To address these limitations, we propose AffordVLA, an affordance-enhanced VLA framework that internalizes manipulation-centric affordance perception into VLA visual representations through implicit representation alignment. Specifically, we construct a zero-shot affordance teacher to extract task-conditioned affordance visual representations from RGB observations and language instructions. AffordVLA aligns the intermediate visual representations of the VLA with the affordance visual representations extracted by the teacher, thereby implicitly injecting manipulation-centric affordance perception into VLA visual representations and improving action accuracy. Extensive simulation and real-world experiments demonstrate that AffordVLA and its affordance teacher achieve state-of-the-art performance and outperform strong baselines. Ablation analyses show that AffordVLA effectively reshapes VLA visual representations while preserving inference efficiency, leading to improved manipulation success rates and training efficiency.

机器人操作功能感知视觉表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。