用信息论重分布动作误差,让机器人更稳准地执行指令。
Reshaping Action Error Distributions for Reliable Vision-Language-Action Models
- 引入最小误差熵机制,优化动作预测的分布而非单点误差。
- 在仿真与真实机器人任务中,成功率显著提升且抗噪声能力增强。
- 无需额外推理开销,适合对鲁棒性要求高的实际应用。
在机器人操作中,视觉-语言-动作(VLA)模型已成为学习通用、可扩展机器人策略的有前景范式。现有大多数VLA框架依赖标准监督目标,如离散动作的交叉熵和连续动作回归的均方误差(MSE),对单个预测施加强约束。本文聚焦连续动作VLA模型,突破传统MSE回归,通过训练过程中重塑动作误差分布来改进性能。基于信息论原理,将最小误差熵(MEE)引入现代VLA架构,提出一种轨迹级MEE目标及其两种加权变体,与MSE联合用于连续动作VLA训练。我们在多个代表性VLA架构上,使用LIBERO和SimplerEnv等仿真基准以及真实机器人操作任务,评估了方法在标准、少样本和噪声环境下的表现。实验结果表明,该方法在各类设置下均实现成功率和鲁棒性的持续提升。在数据不平衡情况下,性能增益仍保持在明确的操作范围内,且训练成本几乎不变,不影响推理效率。我们还提供了理论分析,解释MEE监督的有效性并刻画其实际适用范围。
原文摘要 · Abstract (English)
In robotic manipulation, vision-language-action (VLA) models have emerged as a promising paradigm for learning generalizable and scalable robot policies. Most existing VLA frameworks rely on standard supervised objectives, typically cross-entropy for discrete actions and mean squared error (MSE) for continuous action regression, which impose strong pointwise constraints on individual predictions. In this work, we focus on continuous-action VLA models and move beyond conventional MSE-based regression by reshaping action error distributions during training. Drawing on information-theoretic principles, we introduce Minimum Error Entropy (MEE) into modern VLA architectures and propose a trajectory-level MEE objective, together with two weighted variants, combined with MSE for continuous-action VLA training. We evaluate our approaches across standard, few-shot, and noisy settings on multiple representative VLA architectures, using simulation benchmarks such as LIBERO and SimplerEnv as well as real-world robotic manipulation tasks. Experimental results demonstrate consistent improvements in success rates and robustness across these settings. Under imbalanced data regimes, the gains persist within a well-characterized operating range, while incurring negligible additional training cost and no impact on inference efficiency. We further provide theoretical analyses that explain why MEE-based supervision is effective and characterize its practical range. Project Page: https://cognition2actionlab.github.io/VLA-TMEE.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。