提出一种无需显式建模的单步动作生成方法,提升机器人控制精度。
Implicit Drifting Policy: One-Step Action Generation via Conditional Expert Geometry

- 通过隐式提取专家动作的局部几何结构,实现条件约束建模
- 在单步生成中保持动作流形有效性,实测优于显式漂移方法
- 适用于高频率机器人操控,尤其适合真实场景任务
基于扩散或流匹配的生成式动作策略在行为克隆中表现优异,但其迭代采样难以用于高频机器人控制。尽管近期单步方法缓解了延迟问题,却丢失了训练时重要的动作修正过程。直接通过显式估计训练期漂移场来恢复该机制在数学上是病态的,因示范数据极度稀疏。本文提出隐式漂移策略(IDP),一种单步模仿学习框架,无需显式向量场估计即可将训练时的漂移修正引入策略学习。IDP 从观察相似的专家动作局部变化中提取条件专家几何,并与全局参考几何对比,分离出条件特异性约束,进而自适应加权标量势能目标。结合专家邻近的终端评估,IDP 在训练中直接对单步生成器施加流形约束。在 2D、3D 及真实世界操作任务中的大量实验表明,IDP 有效维持了合法动作流形的遵循性,优于显式漂移方法,并达到与强基线相当的性能。
原文摘要 · Abstract (English)
Generative action policies based on diffusion or flow matching excel in behavior cloning, yet their iterative sampling is prohibitive for high-frequency robot control. While recent one-step formulations alleviate this latency, they inevitably discard the intermediate trajectory evolution that provides crucial action correction. Directly recovering this mechanism by explicitly estimating a training-time drifting field is mathematically ill-posed due to extreme conditional demonstration sparsity. We introduce Implicit Drifting Policy (IDP), a one-step imitation learning framework that brings the training-time correction of Drifting into policy learning without explicit vector field estimation. IDP extracts a conditional expert geometry from the local variation of observation-similar expert actions, and compares it against a global reference geometry to isolate condition-specific constraints. This local geometric structure adaptively weights a scalar potential objective. Combined with an expert-proximal terminal evaluation, IDP directly enforces manifold constraints on the one-step generator during training. Extensive evaluations across 2D, 3D, and real-world manipulation tasks show IDP effectively maintains adherence to valid action manifolds, improving upon explicit drifting methods and achieving competitive performance with strong one-step baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。