arXiv:2605.12369cs.RO2026-05中稿 · RSS 2026被引 5

通过显式引导动作解码,提升机器人模型的泛化能力

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization

论文配图:GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization
图 1 · 摘自论文原文
  • 将动作解码器拆分为多个专用注意力头,分别聚焦物体定位、空间几何和时间技能逻辑
  • 在仿真与真实机器人任务中,成功率显著优于现有基线模型,跨域表现更优
  • 可解释性强,各模块特征解耦且与任务性能正相关,适合需要可靠决策的机器人系统

视觉-语言-动作(VLA)模型通过将动作作为视觉-语言模型中的一个模态,实现通用机器人学习。现有VLA依赖端到端监督隐式学习任务相关特征,但缺乏显式引导时易受视觉捷径或环境噪声干扰,导致泛化能力受限。本文提出GuidedVLA框架,通过手动定义辅助信号,显式引导动作解码器关注任务相关因素。核心思想是将动作解码器视为功能组件的集合,每个注意力头独立训练以捕捉特定因子:物体定位、空间几何和时间技能逻辑。在仿真与真实机器人实验中,GuidedVLA在领域内和跨域设置下均显著提升成功率。此外,这些专化因子的质量与任务表现正相关,且生成特征解耦清晰。结果表明,显式引导动作解码学习是构建更鲁棒、通用的VLA模型的有效方向。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models aim for general robot learning by aligning action as a modality within powerful Vision-Language Models (VLMs). Existing VLAs rely on end-to-end supervision to implicitly enable the action decoding process to learn task-relevant features. However, without explicit guidance, these models often overfit to spurious correlations, such as visual shortcuts or environmental noise, limiting their generalization. In this paper, we introduce GuidedVLA, a framework designed to manually guide the action generation to focus on task-relevant factors. Our core insight is to treat the action decoder not as a monolithic learner, but as an assembly of functional components. Individual attention heads are supervised by manually defined auxiliary signals to capture distinct factors. As an initial study, we instantiate this paradigm with three specialized heads: object grounding, spatial geometry, and temporal skill logic. Across simulation and real-robot experiments, GuidedVLA improves success rates in both in-domain and out-of-domain settings compared to strong VLA baselines. Finally, we show that the quality of these specialized factors correlates positively with task performance and that our mechanism yields decoupled, high-quality features. Our results suggest that explicitly guiding action-decoder learning is a promising direction for building more robust and general VLA models.

机器人学习视觉语言动作注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。