arXiv:2605.22671cs.CV2026-05被引 6

提出新框架让机器人在不同环境下更稳定地执行复杂任务。

From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model

论文配图:From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model
图 1 · 摘自论文原文
  • 用因果Mamba结构捕捉长时间动作轨迹,生成统一行为表示。
  • 实测在三个数据集上表现领先,真实场景迁移仅需一半演示数据。
  • 适合需要高效学习和强泛化能力的机器人控制研究者。

视觉-语言-动作(VLA)模型在分布偏移下性能易下降,因难以跨环境学习通用行为表示。现有方法依赖以动作为中心的潜在变量,但受限于短时序碎片化与静态执行对齐,导致复杂场景下行为不一致。为此,本文提出BehaviorVLA框架,通过学习时序连贯的行为表示实现鲁棒操作。该框架包含两个对称组件:(1) 视动行为编码器(VBE),采用因果Mamba架构聚合长时序轨迹信息,形成统一行为表示;(2) 阶段条件行为解码器(PBD),通过动态对齐任务先验与实时执行进度,精准生成动作。在RoboTwin 2.0、LIBERO和CALVIN上的实验显示,成功率达58%、98%和4.36(平均长度)。值得注意的是,在真实世界仿真到现实的迁移中,BehaviorVLA仅用50%演示数据即达到OpenVLA-OFT的性能,展现优异的数据效率与泛化能力。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models often suffer from performance degradation under distribution shifts, as they struggle to learn generalized behavior representations across varying environments. While existing approaches attempt to construct behavior representations through action-centric latent variables, they are often limited by short-horizon temporal fragmentation and static execution-alignment, leading to inconsistent behaviors in complex scenarios. To address these limitations, we propose \textbf{BehaviorVLA}, a framework that facilitates robust manipulation through the learning of a temporally coherent behavioral representations. Our approach features two symmetric components: (1) the \textbf{Visuomotor Behavior Encoder (VBE)}, which utilizes a causal Mamba-based architecture to aggregate long-horizon trajectory information into a unified behavior representation; and (2) the \textbf{Phase-conditioned Behavior Decoder (PBD)}, which decodes this representation into precise actions by dynamically aligning task-level priors with real-time execution progress. Experiments on RoboTwin 2.0, LIBERO, and CALVIN demonstrate state-of-the-art success rates of 58\%, 98\%, and 4.36 (Avg.Len), respectively. Notably, in real-world sim-to-real transfer, BehaviorVLA matches the performance of OpenVLA-OFT using only 50\% of the demonstration data, showcasing its superior data efficiency and generalization.

机器人控制行为表示时序建模数据效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。