arXiv:2605.00078cs.ROcs.CV2026-05被引 52

用隐变量推理未来,让机器人在不生成画面的情况下提前规划动作。

Being-H0.7: A Latent World-Action Model from Egocentric Videos

  • 引入可学习的隐变量作为感知与动作间的推理接口,实现未来感知。
  • 在6个仿真和多个真实任务中表现最优,无需生成未来画面。
  • 兼顾世界模型的预测能力与直接策略的部署效率,适合实际机器人应用。

视觉-语言-动作模型(VLAs)通过将多模态观测和语言指令直接映射到动作,推动了通用机器人控制的发展,但稀疏的动作监督常导致模型依赖捷径映射而非动态、接触和任务进展的真实表征。近期的世界-动作模型通过视频回放引入未来预测,但像素空间预测成本高且间接,可能建模无关视觉细节并增加训练或推理开销。我们提出 Being-H0.7,一种隐变量世界-动作模型,在不生成未来帧的前提下,将未来感知融入 VLA 式策略中。该模型在感知与动作之间插入可学习的隐变量查询,作为紧凑的推理接口,并采用未来感知的双分支训练设计:可部署的先验分支从当前上下文推断隐状态,而仅用于训练的后验分支则用未来观测嵌入替换这些查询。在隐式推理空间中联合对齐两个分支,使先验分支仅凭当前观测就能推断出未来感知、对动作有用的信息结构。推理时,模型丢弃后验分支,无需视觉回放。在六个仿真基准和多种真实任务上的实验表明,Being-H0.7 达到当前最优或相当性能,兼具世界模型的预测优势与直接 VLA 策略的高效性和可部署性。

原文摘要 · Abstract (English)

Visual-Language-Action models (VLAs) have advanced generalist robot control by mapping multimodal observations and language instructions directly to actions, but sparse action supervision often encourages shortcut mappings rather than representations of dynamics, contact, and task progress. Recent world-action models introduce future prediction through video rollouts, yet pixel-space prediction is a costly and indirect substrate for control, as it may model visual details irrelevant to action generation and introduces substantial training or inference overhead. We present Being-H0.7, a latent world-action model that brings future-aware reasoning into VLA-style policies without generating future frames. Being-H0.7 inserts learnable latent queries between perception and action as a compact reasoning interface, and trains them with a future-informed dual-branch design: a deployable prior branch infers latent states from the current context, while a training-only posterior branch replaces the queries with embeddings from future observations. Jointly aligning the two branches at the latent reasoning space leads the prior branch to reason future-aware, action-useful structure from current observations alone. At inference, Being-H0.7 discards the posterior branch and performs no visual rollout. Experiments across six simulation benchmarks and diverse real-world tasks show that Being-H0.7 achieves state-of-the-art or comparable performance, combining the predictive benefits of world models with the efficiency and deployability of direct VLA policies.

机器人控制隐变量模型未来预测动作规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。