让机器人在有延迟时仍能准确执行动作,提升动态任务鲁棒性。
Delay-Aware Diffusion Policy: Bridging the Observation-Execution Gap in Dynamic Tasks
- 训练和推理时显式建模推理延迟,修正状态偏差。
- 在多种机器人和延迟条件下,成功率显著优于无延迟感知方法。
- 适用于各类策略网络,推动按实际延迟评估模型性能。
机器人感知与执行之间存在数十至数百毫秒的延迟,导致观测状态与执行时刻状态不一致。本文提出延迟感知扩散策略(DA-DP),通过在训练和推理中显式引入推理延迟,将零延迟轨迹校正为延迟补偿轨迹,并以延迟作为条件增强策略。我们在多种任务、机器人及延迟场景下验证了该方法,结果表明其成功率对延迟变化更具鲁棒性。该框架与架构无关,可推广至非扩散策略,为延迟感知模仿学习提供通用范式。更广泛地,推动评估协议从仅报告任务难度转向以实测延迟为基准进行性能分析。
原文摘要 · Abstract (English)
As a robot senses and selects actions, the world keeps changing. This inference delay creates a gap of tens to hundreds of milliseconds between the observed state and the state at execution. In this work, we take the natural generalization from zero delay to measured delay during training and inference. We introduce Delay-Aware Diffusion Policy (DA-DP), a framework for explicitly incorporating inference delays into policy learning. DA-DP corrects zero-delay trajectories to their delay-compensated counterparts, and augments the policy with delay conditioning. We empirically validate DA-DP on a variety of tasks, robots, and delays and find its success rate more robust to delay than delay-unaware methods. DA-DP is architecture agnostic and transfers beyond diffusion policies, offering a general pattern for delay-aware imitation learning. More broadly, DA-DP encourages evaluation protocols that report performance as a function of measured latency, not just task difficulty.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。