arXiv:2605.27947cs.RO2026-05被引 2

动态调整视频去噪程度,让机器人决策更准更快。

SANTS: A State-Adaptive Scheduler for World Action Models

论文配图:SANTS: A State-Adaptive Scheduler for World Action Models
图 1 · 摘自论文原文
  • 根据当前视频状态自适应选择去噪深度,而非固定终点。
  • 在真实机器人任务中达73.1%成功率,延迟降低超79%。
  • 适合需要高效视觉决策的机器人系统开发者。

世界动作模型(WAMs)通过基于视频的未来表征来提升机器人操作能力。然而,在像素空间的WAMs中,最佳动作条件并不总是完全去噪的视频。控制性去噪深度扫描表明,视频优化可减少动作误差至某一状态相关点,之后增益可能饱和甚至逆转,因后期预测变得不具动作相关性或物理可靠性。这提示动作生成应沿视频噪声轨迹选择状态相关的停止点,而非固定终点。我们提出状态自适应噪声轨迹调度器(SANTS),一种轻量级视频到动作扩散策略的调度器。在每个视频决策点,SANTS读取当前视频状态表示与噪声水平,联合预测累积停止危险和相对噪声进展比例。SANTS通过冻结动作分支生成最终动作块后计算的路径级奖励进行后训练,因此调度器被优化为提升下游动作质量,而非中间视频保真度,同时显式惩罚冗余视频状态更新。实验显示,SANTS在RoboTwin 2.0上达到94.4%总体成功率,在七个真实机器人任务上平均成功率73.1%,分别比完整去噪延迟降低81.7%和79.0%。结果表明,沿视频噪声轨迹自适应选择可保留WAM式未来推理的控制优势,同时消除大量冗余推理开销。

原文摘要 · Abstract (English)

World Action Models (WAMs) improve robot manipulation by using video-based future representations to condition action generation. In pixel-space WAMs, however, the best action condition is not necessarily the fully denoised video. Controlled denoising-depth scans show that video refinement can reduce action error up to a state-dependent point, after which the gain may saturate or even reverse when late predictions become less action-relevant or physically unreliable. This suggests that action generation should use a state-dependent point along the video noise trajectory rather than a fixed terminal denoising depth. We introduce State-Adaptive Noise Trajectory Scheduler (SANTS), a lightweight scheduler for video-to-action diffusion policies. At each video decision point, SANTS reads the current video-state representation and noise level, then jointly predicts a cumulative stopping hazard and a relative noise-progression ratio. SANTS is post-trained with a path-level reward computed after the frozen action branch generates the final action chunk, so the scheduler is optimized for downstream action quality rather than intermediate video fidelity, while redundant video-state updates are explicitly penalized. Experiments show that SANTS reaches \(94.4\%\) overall success on RoboTwin 2.0 and \(73.1\%\) average success across seven real-robot tasks, while reducing latency by \(81.7\%\) and \(79.0\%\) relative to full video denoising, respectively. These results indicate that adaptive selection along the video noise trajectory can preserve the control benefits of WAM-style future reasoning while removing much of its redundant inference cost.

机器人去噪调度扩散模型动作生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。