arXiv:2607.15714cs.RO2026-07被引 2

提升机器人在陌生组合任务中的泛化能力,解决视觉依赖与动作过拟合问题。

AC-VLA: Robust Out-of-Distribution Action Execution via Compositional Learning

论文配图:AC-VLA: Robust Out-of-Distribution Action Execution via Compositional Learning
图 1 · 摘自论文原文
  • 通过分解指令与姿态对齐生成子任务监督信号,实现组合式学习
  • 在LIBERO-OOD上提升28%的陌生任务成功率,保持原性能不变
  • 无需修改模型结构,可适配任意视觉-语言-动作框架

视觉-语言-动作(VLA)模型在端到端机器人操作中表现优异,但在熟悉子任务以未见方式重组时,面临分布外(OOD)泛化难题。我们识别出两种相互强化的失败模式:轨迹过拟合,即模型过度依赖整体轨迹模式而非组合式子技能语义;感知捷径,即动作标记过度依赖腕部视角纹理,忽视全局空间定位。为此,我们提出AC-VLA——一个无需架构修改的组合式学习框架,包含两个通用组件:(i) 使用大语言模型驱动的指令分解器与本体感知轨迹对齐器生成密集子任务监督信号,并在完整演示与分解数据上混合训练,赋予模型组合泛化能力;(ii) 状态条件下的非对称掩码策略,在闭爪阶段抑制腕部视角输入,强制全局语义定位。所有组件均可直接嵌入任意VLA主干网络。在π₀.₅上实例化,并在LIBERO与LIBERO-OOD基准上评估,AC-VLA在组合式分布外任务上实现约28%的绝对性能提升,同时维持近乎完美的分布内表现。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models excel at end-to-end robotic manipulation but struggle with out-of-distribution (OOD) generalization when familiar sub-tasks are recombined in unseen configurations. We identify two mutually reinforcing failure modes: \emph{trajectory overfitting}, where models overfit to holistic trajectory patterns rather than compositional sub-skill semantics; and \emph{perceptual shortcut}, where action tokens over-rely on wrist-view textures at the expense of global spatial grounding. To address both, we introduce \textbf{AC-VLA}, a plug-and-play Action Compositional learning framework comprising two architecture-agnostic components: \textbf{(i)} a compositional learning module that uses an LLM-driven instruction decomposer and a proprioceptive trajectory aligner to generate dense sub-task supervision, followed by mixed training on complete demonstrations and decomposed data to endow the model with compositional generalization; and \textbf{(ii)} a state-conditioned asymmetric masking strategy that suppresses wrist-view inputs during closed-gripper phases, enforcing global semantic grounding. All components are architectural modification-free and directly integrable into any VLA backbone. Instantiated on $π_{0.5}$ and evaluated on LIBERO and LIBERO-OOD benchmarks, AC-VLA achieves a ~28% absolute improvement on compositional OOD tasks while maintaining near-perfect in-distribution performance.

机器人控制组合泛化视觉语言动作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。