让机器人动作与视频生成一步完成,速度提升23倍且保持高成功率。
Flash-WAM: Modality-Aware Distillation for World Action Models

- 分模态设计噪声适配的压缩方法,分别处理动作与视频流的差异噪声分布。
- 在仿真和真实机器人上均实现95%以上成功率,单步推理延迟降至348毫秒。
- 适合需要实时控制的具身智能系统,尤其对硬件资源有限的场景友好。
世界动作模型(WAMs)通过迭代扩散联合生成未来视频与机器人动作,在操作基准测试中表现优异,但需数十次去噪步骤,难以支持实时控制。已有步骤蒸馏方法在联合视频-动作场景中失效,因视频与动作流采用不同信噪比偏移的噪声调度,导致训练时边际噪声分布显著不同。本文提出 extbf{Flash-WAM},一种受一致性蒸馏启发的模态感知蒸馏框架:为动作流(低噪声域)设计线性梯度缩放参数化,为视频流(高噪声域)采用方差保持参数化,基于一致性函数族的结构分析,确保梯度缩放满足一致性边界条件。该方法在LingBot-VA上实现每模态单步推理。在RoboTwin 2.0上,推理延迟从8.1秒降至348毫秒(NVIDIA L40S),提速23倍,支持实时推理。性能方面,仿真任务成功率保持在85.5%(RoboTwin 2.0)、95.7%(LIBERO),真实机器人平均成功率恢复至60%(Unitree G1人形机器人),而朴素一致性蒸馏仅剩24%。
原文摘要 · Abstract (English)
World-action models (WAMs) jointly generate future video and robot actions through iterative diffusion, achieving strong performance on manipulation benchmarks but requiring tens of denoising steps, a cost that precludes real-time control. Step distillation has emerged as the natural remedy, but off-the-shelf methods break down in the joint video-action setting because video and action streams use different SNR-shifted noise schedules and reach training with substantially different marginal noise distributions, an asymmetry that single-modality distillation methods cannot accommodate. We introduce \textbf{Flash-WAM}, a modality-aware step-distillation framework inspired by consistency distillation that selects the consistency function for each modality to match its noise regime: a linear-gradient-scaling parametrization for the action stream's low-noise regime, paired with a variance-preserving parametrization for the video stream's high-noise regime, grounded in a structural analysis of the consistency-function family that characterizes the achievable gradient scaling under the consistency boundary condition. Instantiated on LingBot-VA, Flash-WAM compresses inference to a single step in each modality. On RoboTwin 2.0, this reduces per-chunk latency from $8.1$ seconds to $348$ ms on NVIDIA L40S, a $23{\times}$ speedup that enables real-time inference. Flash-WAM preserves task success on simulation benchmarks ($85.5\%$ RoboTwin 2.0, $95.7\%$ LIBERO) and substantially recovers real-world performance ($60\%$ average on a Unitree G1 humanoid robot), while naive consistency distillation drops to $24\%$ at the same step budget.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。