arXiv:2604.09824cs.ROcs.CL2026-04

让机器人听懂指令并识别模糊信息,提升视觉语言动作模型的鲁棒性。

ProGAL-VLA: Grounded Alignment through Prospective Reasoning in Vision-Language-Action Models

论文配图:ProGAL-VLA: Grounded Alignment through Prospective Reasoning in Vision-Language-Action Models
图 1 · 摘自论文原文
  • 构建3D实体中心图谱,用符号化目标与视觉实体对齐
  • 在LIBERO-Plus上语言忽略减少3-4倍,准确率从30.3%升至71.5%
  • 能主动识别模糊指令并请求澄清,适合复杂交互场景

视觉语言动作(VLA)模型使通用机器人成为可能,但常出现语言忽视问题,依赖视觉捷径且对指令变化不敏感。本文提出基于前瞻性推理的接地对齐模型(ProGAL-VLA),构建3D实体中心图谱(GSM),通过慢速规划生成符号化子目标,并利用接地对齐对比损失(GAC)将其与可感知实体对齐。所有动作均基于验证后的目标嵌入 $g_t$,其注意力熵提供内在歧义信号。在LIBERO-Plus上,模型在机器人扰动下鲁棒性从30.3%提升至71.5%,语言忽视降低3-4倍,实体检索召回率@1从0.41提升至0.71。在自定义歧义基准测试中,达到AUROC 0.81(对比基线0.52)、AUPR 0.79,模糊输入澄清率从0.09升至0.81,同时未损害非模糊任务成功率。验证瓶颈提升了语言-动作互信息,GAC损失施加了实体级InfoNCE约束,注意力熵实现校准的选择性预测,表明显式验证接地是实现指令敏感、歧义感知智能体的有效路径。

原文摘要 · Abstract (English)

Vision language action (VLA) models enable generalist robotic agents but often exhibit language ignorance, relying on visual shortcuts and remaining insensitive to instruction changes. We present Prospective Grounding and Alignment VLA (ProGAL-VLA), which constructs a 3D entity-centric graph (GSM), uses a slow planner to produce symbolic sub-goals, and aligns them with grounded entities via a Grounding Alignment Contrastive (GAC) loss. All actions are conditioned on a verified goal embedding $g_t$, whose attention entropy provides an intrinsic ambiguity signal. On LIBERO-Plus, ProGAL-VLA increases robustness under robot perturbations from 30.3 to 71.5 percent, reduces language ignorance by 3x-4x, and improves entity retrieval from 0.41 to 0.71 Recall@1. On the Custom Ambiguity Benchmark, it reaches AUROC 0.81 (vs. 0.52), AUPR 0.79, and raises clarification on ambiguous inputs from 0.09 to 0.81 without harming unambiguous success. The verification bottleneck increases mutual information of language-actions, the GAC loss imposes an entity-level InfoNCE bound, and attention entropy yields calibrated selective prediction, indicating that explicit verified grounding is an effective path toward instruction-sensitive, ambiguity-aware agents.

机器人多模态指令理解智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。