arXiv:2605.07794cs.RO2026-05被引 2

让每帧潜在表示自适应调整去噪时间,提升机器人动作生成的可靠性。

NoiseGate: Learning Per-Latent Timestep Schedules as Information Gating in World Action Models

论文配图:NoiseGate: Learning Per-Latent Timestep Schedules as Information Gating in World Action Models
图 1 · 摘自论文原文
  • 为每帧潜在变量设计可学习的去噪时间调度策略
  • 在多种机器人操作任务上显著提升动作生成性能
  • 无需人工设定形状先验,适合复杂场景下的自主控制

世界动作模型(WAMs)是一类将机器人动作生成与未来观测建模相结合的新范式。本文聚焦于视频与动作联合建模框架,即通过共享的去噪或流轨迹同步生成动作与想象中的未来视觉,使感知、预测与控制在单一生成过程中耦合。现有WAMs通常采用混合变压器(MoT)架构,通过共享自注意力实现视频与动作标记的交互。尽管该架构理论上可为每帧预测潜在状态分配独立的时间步 $t_f$,但当前系统将其压缩为单一共享标量 $t$。基于扩散强制的噪声即掩码视角,这种统一调度强加了不合理的先验:所有预测潜在状态对动作生成的可靠性相同。本文提出将每帧潜在状态的调度视为可学习的信息门控策略:通过调整潜在帧的噪声水平,动态调节其键/值对动作标记的贡献可靠性。我们提出NoiseGate,包含三个关键组件:骨干训练阶段的独立每潜在变量时间步采样、去噪过程中输出每潜在变量时间增量的轻量级门控网络,以及无需手工形状先验的任务奖励优化策略。基于联合视频-动作MoT骨干,NoiseGate在多样化的RoboTwin随机场景操作任务中均取得一致性能提升。

原文摘要 · Abstract (English)

World Action Models (WAMs) are an emerging family of policies that tie robot action generation to future-observation modeling. In this work, we focus on the joint video--action modeling paradigm, where actions and imagined future observations are co-generated along a shared denoising or flow trajectory, so that perception, prediction, and control are coupled within one generative process. Existing WAMs typically realize this paradigm with a Mixture-of-Transformers (MoT), where video and action tokens interact through shared self-attention. This architecture can in principle assign a separate timestep $t_f$ to each predicted latent frame, yet current systems collapse this degree of freedom onto a single shared scalar $t$. Under the noise-as-masking view of Diffusion Forcing, this shared schedule imposes the unjustified prior that every predicted latent is equally reliable for action generation. We instead view the per-latent schedule as a \emph{learnable information-gating policy}: by changing a latent frame's noise level, the policy modulates the reliability of its Key/Value contribution to the action tokens. We propose \textbf{NoiseGate}, which combines independent per-latent timestep sampling during backbone training, a lightweight Gating Policy Network that emits per-latent time increments during denoising, and task-reward optimization that trains the schedule policy without hand-crafted shape priors. Built on a joint video--action MoT backbone, NoiseGate delivers consistent gains on diverse RoboTwin random-scene manipulation tasks.

动作生成扩散模型机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。