arXiv:2602.21445cs.RO2026-02被引 8

动态调整机器人动作执行范围,提升视觉语言模型适应能力

VLA Knows Its Limits: Adaptive Execution Horizons for Robot Policies

  • 用动作自注意力判断预测极限,动态调整每步执行动作数
  • 实测在仿真与真实任务中均显著提升性能,且计算开销极低
  • 适合追求鲁棒性与泛化能力的机器人决策研究者

基于视觉-语言-动作(VLA)模型的动作分块已成为主流,但执行范围(即每段预测动作中实际执行的数量)的选择尚未深入探索。本文发现,随着执行范围增大,性能先升后降。通过分析流式VLA中的交叉与自注意力权重,揭示两个关键现象:(i) 段内动作对视觉-语言标记的关注保持不变,难以响应环境变化;(ii) 初末动作标记作为稳定锚点,构成中间动作的潜在中心。基于此,我们提出AutoHorizon——首个测试时动态估计执行范围的方法,利用动作自注意力作为模型预测能力的代理。在多种仿真与真实世界机器人操作任务中,该方法表现优异,计算开销可忽略,并能跨任务、跨模型泛化。

原文摘要 · Abstract (English)

Action chunking has recently emerged as a standard practice in flow-based Vision-Language-Action (VLA) models. However, the effect and choice of the execution horizon - the number of actions to be executed from each predicted chunk - remains underexplored. In this work, we first show that varying the execution horizon leads to substantial performance deviations, with performance initially improving and then declining as the horizon increases. To uncover the reasons, we analyze the cross- and self-attention weights in flow-based VLAs and reveal two key phenomena: (i) intra-chunk actions attend invariantly to vision-language tokens, limiting adaptability to environmental changes; and (ii) the initial and terminal action tokens serve as stable anchors, forming latent centers around which intermediate actions are organized. Motivated by these insights, we interpret action self-attention weights as a proxy for the model's predictive limit and propose AutoHorizon, the first test-time method that dynamically estimates the execution horizon for each predicted action chunk to adapt to changing perceptual conditions. Across simulated and real-world robotic manipulation tasks, AutoHorizon is performant, incurs negligible computational overhead, and generalizes across diverse tasks and flow-based models.

机器人控制视觉语言模型动态执行自注意力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。