arXiv:2606.22729cs.RO2026-06中稿 · ICRA被引 2

用世界模型让机器人行为满足时间逻辑约束,不重训也能大幅减少违规。

Temporal Logic Guidance for Action-Only Diffusion Policies with World Models

  • 用独立世界模型计算时序逻辑鲁棒性并反向传播梯度
  • 在机器人任务中将约束违反率从80%降至4%,成功率仍100%
  • 适合需要安全可靠行为的交互式机器人场景

扩散策略能生成多模式机器人行为,但在推理时难以灵活选择行为模式,这在人机协作中是个缺陷。现有方法虽用信号时序逻辑(STL)表达人类意图来引导扩散策略,但仅适用于同时生成动作和状态的复杂模型,导致计算开销大。本文提出一种新方法,针对仅生成动作的扩散策略,通过独立训练的世界模型实现STL鲁棒性的可微评估,并将梯度注入扩散过程,从而在不重训练的情况下引导行为满足约束,提升合规性同时保持任务性能。在Robomimic的Can Transport任务上,该方法维持100%任务成功率,将约束违规率从基线超过80%降至4%。还讨论了提升鲁棒性和处理更复杂约束的扩展方向。

原文摘要 · Abstract (English)

Diffusion policies enable multimodal robot behavior but offer limited ability to choose among behavior modes at inference time, even though such control is desirable in human-robot settings. Prior solutions to this lack of control have utilized Signal Temporal Logic (STL) to express human intentions and provide corresponding guidance for diffusion policy inference. However, these approaches can only guide diffusion policies that jointly generate future actions and states, increasing both complexity and runtime. We propose a novel guidance method for action-only diffusion policies that uses a separate learned world model to enable differentiable evaluation of STL robustness, with its gradient then injected into the diffusion process. This steers behavior toward constraint satisfaction without retraining, improving constraint adherence while preserving task performance. On the Can Transport task from Robomimic, our method maintains 100% task success while reducing constraint violations from over 80% for baseline methods to 4%. We also discuss extensions toward improved robustness and more complex constraints.

扩散模型机器人时序逻辑世界模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。