arXiv:2605.13548cs.ROcs.AI2026-05被引 2

让机器人模型更关注关键动作,提升复杂操作成功率。

AttenA+: Rectifying Action Inequality in Robotic Foundation Models

论文配图:AttenA+: Rectifying Action Inequality in Robotic Foundation Models
图 1 · 摘自论文原文
  • 根据动作速度动态调整训练权重,聚焦关键低速操作段。
  • 在Libero和RoboTwin数据集上分别提升1.5%和0.6%准确率。
  • 无需修改模型结构,可直接嵌入现有机器人模型使用。

现有机器人基础模型依赖时间同质性假设,对所有动作赋予同等重要性,忽视了操作中的物理层级差异。实际上,低速段常决定任务成败,而高速段多为容错过渡。这种损失权重均匀与物理重要性不匹配的问题限制了视觉-语言-动作(VLA)和世界-动作模型(WAM)在复杂长程任务中的表现。为此,我们提出AttenA+,一种与架构无关的框架,通过速度驱动的动作注意力机制,基于逆速度场重加权训练目标,使学习能力自然对齐物理需求。作为即插即用增强模块,AttenA+无需结构修改或额外参数。大量实验表明,其显著提升当前领先模型上限:在Libero基准上将OpenVLA-OFT提升至98.6%(+1.5%),在RoboTwin 2.0上将FastWAM提升至92.4%(+0.6%)。真实机械臂验证进一步证明其鲁棒性和跨任务泛化能力。本工作表明,挖掘动作序列内在结构先验,是标准缩放定律之外高效、物理感知的补充路径,为通用机器人控制开辟新方向。

原文摘要 · Abstract (English)

Existing robotic foundation models, while powerful, are predicated on an implicit assumption of temporal homogeneity: treating all actions as equally informative during optimization. This "flat" training paradigm, inherited from language modeling, remains indifferent to the underlying physical hierarchy of manipulation. In reality, robot trajectories are fundamentally heterogeneous, where low-velocity segments often dictate task success through precision-demanding interactions, while high-velocity motions serve as error-tolerant transitions. Such a misalignment between uniform loss weighting and physical criticality fundamentally limits the performance of current Vision-Language-Action (VLA) models and World-Action Models (WAM) in complex, long-horizon tasks. To rectify this, we introduce AttenA+, an architecture-agnostic framework that prioritizes kinematically critical segments via velocity-driven action attention. By reweighting the training objective based on the inverse velocity field, AttenA+ naturally aligns the model's learning capacity with the physical demands of manipulation. As a plug-and-play enhancement, AttenA+ can be integrated into existing backbones without structural modifications or additional parameters. Extensive experiments demonstrate that AttenA+ significantly elevates the ceilings of current state-of-the-art models. Specifically, it improves OpenVLA-OFT to 98.6% (+1.5%) on the Libero benchmark and pushes FastWAM to 92.4% (+0.6%) on RoboTwin 2.0. Real-world validation on a Franka manipulator further showcases its robustness and cross-task generalization. Our work suggests that mining the intrinsic structural priors of action sequences offers a highly efficient, physics-aware complement to standard scaling laws, paving a new path for general-purpose robotic control.

机器人控制动作优化模型增强物理感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。