通过历史轨迹增强视觉输入,解决长程机器人操作中的动作歧义问题
Trace-Focused Diffusion Policy for Multi-Modal Action Disambiguation in Long-Horizon Robotic Manipulation
- 用执行轨迹显式建模历史,引导当前动作决策
- 在动作歧义任务上性能提升80.56%,抗干扰能力提升86.11%
- 适合复杂场景下的长程机器人控制,计算开销仅增加6.4%
基于生成模型的策略在模仿学习的机器人操作中表现优异,能从示范中学习动作分布。然而,在长程任务中,不同阶段可能出现视觉相似的观察,但需采取不同动作,仅依赖瞬时观察会导致预测歧义,称为多模态动作歧义(MA2)。为此,我们提出轨迹聚焦扩散策略(TF-DP),一种基于扩散的简单而有效的框架,显式地将动作生成条件化于机器人的执行历史。TF-DP将历史运动表示为显式的执行轨迹,并将其投影到视觉观察空间,当当前观察不足以判断时提供阶段感知上下文。此外,诱导的轨迹聚焦场强调与历史运动相关联的任务关键区域,提升了对背景视觉干扰的鲁棒性。我们在真实世界中存在明显多模态动作歧义和视觉杂乱的机器人操作任务上评估了TF-DP。实验结果表明,该方法显著提升了时间一致性与鲁棒性,在多模态动作歧义任务上比基线扩散策略提升80.56%,在视觉干扰下提升86.11%,同时推理效率保持较高,仅增加6.4%的运行时间。这些结果证明,执行轨迹条件化是一种可扩展且原则性的单策略长程机器人操作鲁棒解决方案。
原文摘要 · Abstract (English)
Generative model-based policies have shown strong performance in imitation-based robotic manipulation by learning action distributions from demonstrations. However, in long-horizon tasks, visually similar observations often recur across execution stages while requiring distinct actions, which leads to ambiguous predictions when policies are conditioned only on instantaneous observations, termed multi-modal action ambiguity (MA2). To address this challenge, we propose the Trace-Focused Diffusion Policy (TF-DP), a simple yet effective diffusion-based framework that explicitly conditions action generation on the robot's execution history. TF-DP represents historical motion as an explicit execution trace and projects it into the visual observation space, providing stage-aware context when current observations alone are insufficient. In addition, the induced trace-focused field emphasizes task-relevant regions associated with historical motion, improving robustness to background visual disturbances. We evaluate TF-DP on real-world robotic manipulation tasks exhibiting pronounced multi-modal action ambiguity and visually cluttered conditions. Experimental results show that TF-DP improves temporal consistency and robustness, outperforming the vanilla diffusion policy by 80.56 percent on tasks with multi-modal action ambiguity and by 86.11 percent under visual disturbances, while maintaining inference efficiency with only a 6.4 percent runtime increase. These results demonstrate that execution-trace conditioning offers a scalable and principled approach for robust long-horizon robotic manipulation within a single policy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。