arXiv:2603.11984cs.CV2026-03被引 2

让机器人用单步推理实现多模式精准操作,速度提升10倍

Ada3Drift: Adaptive Training-Time Drifting for One-Step 3D Visuomotor Robotic Manipulation

  • 把迭代优化从推理阶段移到训练阶段,用漂移场引导动作向专家示范聚集
  • 单步生成(1 NFE)即达顶尖性能,函数评估次数减少10倍
  • 适合少样本、高实时性要求的3D机械臂操控任务

基于扩散模型的视觉-运动策略通过迭代去噪有效捕捉多模态动作分布,但高推理延迟限制了实时机器人控制。近期流匹配与一致性方法实现单步生成,却牺牲了多模态保真度,将多样行为压缩为平均化且常不物理可行的轨迹。我们观察到机器人领域算力分配不对称(离线训练 vs. 实时推理),自然启发将迭代优化从推理迁移至训练阶段以恢复多模态精度。基于此,提出Ada3Drift,学习一个训练时的漂移场,使预测动作被吸引至专家示范模式,同时排斥远离其他生成样本,从而实现从3D点云观测的高保真单步生成(1 NFE)。为应对少样本场景,引入符号函数调度损失,实现粗粒度分布学习到模式锐化优化的渐进过渡,并采用多尺度场聚合以捕捉不同空间粒度的动作模式。在三个仿真基准(Adroit、Meta-World、RoboTwin)及真实机器人操控任务上的实验表明,Ada3Drift达到当前最优性能,相比扩散基方法减少10倍函数评估次数。

原文摘要 · Abstract (English)

Diffusion-based visuomotor policies effectively capture multimodal action distributions through iterative denoising, but their high inference latency limits real-time robotic control. Recent flow matching and consistency-based methods achieve single-step generation, yet sacrifice the ability to preserve distinct action modes, collapsing multimodal behaviors into averaged, often physically infeasible trajectories. We observe that the compute budget asymmetry in robotics (offline training vs.\ real-time inference) naturally motivates recovering this multimodal fidelity by shifting iterative refinement from inference time to training time. Building on this insight, we propose Ada3Drift, which learns a training-time drifting field that attracts predicted actions toward expert demonstration modes while repelling them from other generated samples, enabling high-fidelity single-step generation (1 NFE) from 3D point cloud observations. To handle the few-shot robotic regime, Ada3Drift further introduces a sigmoid-scheduled loss transition from coarse distribution learning to mode-sharpening refinement, and multi-scale field aggregation that captures action modes at varying spatial granularities. Experiments on three simulation benchmarks (Adroit, Meta-World, and RoboTwin) and real-world robotic manipulation tasks demonstrate that Ada3Drift achieves state-of-the-art performance while requiring $10\times$ fewer function evaluations than diffusion-based alternatives.

3D操控单步生成多模态机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。