arXiv:2607.28405cs.AIcs.LG2026-07

针对世界动作模型的量化难题,提出精细化校准方法。

QuantWAMs: Calibrating at the Right Granularity for World Action Models

论文配图:QuantWAMs: Calibrating at the Right Granularity for World Action Models
图 1 · 摘自论文原文
  • 按模块坐标兼容性聚合激活证据,实现精准量化
  • 在真实部署环境下保持精度,误差仅0.2–0.7个百分点
  • 适合需高效部署的机器人视觉动作系统

世界动作模型(WAMs)联合预测未来观测与动作,但其迭代去噪和闭环执行导致部署成本高昂。现有训练后量化(PTQ)方法因依赖开环目标、同质模型假设及非部署分布校准,难以适配WAMs。本文提出QuantWAMs,一种将量化决策与模型结构、推演分布和任务目标一致的PTQ框架。引入三项策略:共享基线异常值校准,仅在坐标兼容模块间聚合激活证据;联合目标显著性,基于视频-动作联合梯度计算经验费雪得分,并在稳定层粒度分配权重精度;固定干预推演审计,利用可达闭环状态调整去噪步骤保护策略,不增加精度预算。在Fast-WAM与LingBot-VA上评估,涵盖RoboTwin 2.0、LIBERO及真实机器人操控(AgiBot G2)。W4A4主导设置下,仿真平均性能与FP16相差仅0.2–0.7个百分点。真实机器人实验验证了部署可行性。针对目标视频与动作模块,量化后峰值权重与激活内存降至FP16的约29%,块级速度提升1.4–1.6倍。

原文摘要 · Abstract (English)

World Action Models (WAMs) jointly predict future observations and actions, but their iterative denoising and closed-loop execution make efficient deployment costly. Existing post-training quantization (PTQ) methods are poorly suited to WAMs because they rely on open-loop objectives, homogeneous model assumptions, and calibration distributions that do not reflect deployment. We present QuantWAMs, a PTQ framework that aligns quantization decisions with the calibration context defined by model structure, rollout distribution, and task objective. QuantWAMs introduces three strategies: shared-basis outlier calibration, which pools activation evidence only across coordinate-compatible modules; co-training-objective saliency, which computes empirical-Fisher scores from the joint video--action gradient and assigns weight precision at a calibration-stable layer granularity; and fixed-intervention rollout auditing, which revises denoising-step protection schedules using reachable closed-loop states without changing the precision budget. We evaluate QuantWAMs on Fast-WAM and LingBot-VA across RoboTwin 2.0, LIBERO, and real-robot manipulation with an AgiBot G2. Under a W4A4-dominant setting, the reported simulation means differ from FP16 by 0.2--0.7 percentage points. Real-robot trials further establish deployment feasibility on three manipulation tasks. For the targeted video and action blocks, QuantWAMs reduces peak weight-and-activation memory to about 29\% of FP16 and provides 1.4--1.6$\times$ block-level speedups.

量化机器人动作建模部署优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。