用测试时的原始动作引导,让机器人更智能地完成复杂操作。
PriGo: Test-Time Primitive Guidance to Diffusion and Flow Policies for Adaptive Robotic Manipulation

- 通过轻量级网络实时预测动作原语,指导推理过程
- 在多个数据集和真实机器人上显著提升长期执行成功率
- 无需重训练,适配现有扩散与流模型,适合部署于实际场景
模仿学习推动了机器人操作的显著进展,尤其是基于扩散和流模型的策略能直接从示范中生成复杂的视觉-运动行为。然而,这些策略在任务和环境间泛化能力仍不足,主要原因在于它们倾向于模仿表面的动作相关性,而非底层意图。受人类行为组合结构启发,我们提出PriGo:一种用于鲁棒机器人操作的测试时原始动作引导自适应框架。PriGo引入PANet,一个轻量级原始动作预测模块,可直接从观测中推断原始动作分布。我们进一步提出可微分的原始动作引导机制,在推理过程中优化生成动作,使轨迹趋向语义一致的行为。与以往原始动作条件化方法不同,PriGo完全在测试时运行,无需重新训练,可无缝集成到预训练的扩散与流策略中。在LIBERO、CALVIN、SIMPLER及真实机器人任务上的大量实验表明,PriGo在扩散与流基策略上均持续提升了鲁棒性、长程执行能力与泛化性能。
原文摘要 · Abstract (English)
Imitation learning has enabled remarkable progress in robotic manipulation, especially with diffusion and flow-based policies that generate complex visuomotor behaviors directly from demonstrations. Yet, despite their strong performance, these policies often fail to generalize across tasks and environments. A key reason is that existing policies tend to imitate superficial action correlations rather than the underlying intent. Inspired by the compositional structure of human behaviors, we propose PriGo, a primitive-guided test-time adaptive framework for robust robotic manipulation. PriGo introduces PANet, a lightweight primitive prediction module that infers primitive distributions directly from observations. We further propose a differentiable primitive guidance mechanism that refines generated actions during inference, steering trajectories toward semantically consistent behaviors. Unlike prior primitive-conditioned approaches, PriGo operates entirely at test time and can be seamlessly integrated into pretrained diffusion and flow policies without retraining. Extensive experiments on LIBERO, CALVIN, SIMPLER, and real-world robotic tasks demonstrate that PriGo consistently improves robustness, long-horizon execution, and generalization ability across both diffusion and flow-based policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。