arXiv:2607.11027cs.RO2026-07

用分段扩散模型实现机器人长程自适应操作,兼顾精度与实时性。

SegDiff: Segmented Trajectory Diffusion for Consistent and Adaptive Robot Manipulation

论文配图:SegDiff: Segmented Trajectory Diffusion for Consistent and Adaptive Robot Manipulation
图 1 · 摘自论文原文
  • 将动作分解为关键帧间分段,用扩散模型预测连续轨迹
  • 在模拟和真实场景中均超越现有方法,支持长时序推理
  • 适合需要高精度与实时响应的复杂机械臂任务

模仿学习通过观察-动作映射使机器人从示范中习得操作技能。现有方法或预测短时程连续动作序列,或输出离散关键帧。前者因预测窗口短导致误差累积,难以处理多模态动作分布;后者需依赖外部规划器,限制了实时应用。为此,我们提出SegDiff,一种闭环视觉-运动策略,融合两种范式优势。SegDiff将示范分解为关键帧之间的运动片段,学习从当前状态到下一关键帧的连续轨迹,实现长时程预测并支持实时修正。此外,利用扩散模型与DDIM反演能力,提出动态时间集成机制,使策略能高效响应动态环境,缓解不一致多模态采样带来的不连续问题。SegDiff在多种模拟与真实场景中显著优于现有方法,展现出强时间依赖推理能力,同时保持实时适应性与控制稳定性。

原文摘要 · Abstract (English)

Imitation learning enables robots to acquire manipulation skills from demonstrations by mapping observations to actions. Existing approaches predict either short-horizon continuous action sequences or discrete keyposes. However, continuous prediction methods suffer from compounding errors due to short prediction horizons and struggle with multi-modal action distributions, whereas keypose-based methods necessitate an external planner, constraining real-time applicability. To address these challenges, we introduce SegDiff, a closed-loop visuomotor policy that integrates the strengths of both paradigms. SegDiff decomposes demonstrations into motion segments between keyposes and learns to predict the continuous trajectory from the current state to the next keypose, enabling long-horizon prediction with real-time refinement. Furthermore, we leverage the capability of diffusion models and DDIM inversion to propose a Dynamic Temporal Ensembling mechanism, which allows the policy to efficiently respond to dynamic environments and mitigate discontinuities caused by inconsistent multi-modal sampling. SegDiff demonstrates significant performance gains over existing approaches across various simulated and real-world scenarios, indicating its strong ability to reason over extended temporal dependencies while maintaining real-time adaptability and control stability.

机器人操作扩散模型模仿学习轨迹生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。