arXiv:2606.00837cs.ROcs.LG2026-06

分阶段生成长序列,既保全局合理又提局部细节质量。

Coarse-to-Fine Compositional Diffusion for Long-Horizon Planning

论文配图:Coarse-to-Fine Compositional Diffusion for Long-Horizon Planning
图 1 · 摘自论文原文
  • 先建粗略全局框架,再细化局部结构
  • 相比旧方法提升全局连贯性,样本质量更高
  • 推理效率提高2-8倍,适合长序列生成任务

扩散模型在生成结构化数据方面表现优异,但多数任务需要超出其典型训练规模的输出。组合生成通过整合预训练的短时程先验中的重叠局部规划来实现长时程输出。然而,标准组合方法仅强制相邻局部规划之间的一致性,缺乏对整体结构的直接约束,导致局部兼容的规划仍可能形成不合理的路径、任务序列或时间演化。现有方法通过重复传播局部一致性信号或引入推理时优化提升全局连贯性,但随着局部规划数量或维度增加,计算开销显著上升。我们提出粗到精组合扩散(CoFi),一种推理时采样器,将全局结构构建与局部细节精炼分离。CoFi 首先在共享粗略结构周围对局部去噪估计进行对齐,生成捕捉长程任务级布局的全局骨架;随后将该骨架扩散至中间噪声水平,并使用相同的预训练局部先验进行去噪,恢复局部精细结构的同时保留骨架诱导的全局连贯性。在长时程机器人规划、全景图像生成和长视频生成任务中,CoFi 不仅在全局连贯性和局部样本质量上优于现有组合基线,且所需去噪器评估次数减少2-8倍。

原文摘要 · Abstract (English)

Diffusion models provide strong priors for generating structured data, but many tasks require outputs beyond the scale on which these models are typically trained. Compositional generation addresses this by composing overlapping local plans from a pretrained short-horizon prior into a long-horizon output. However, standard composition primarily enforces agreement between neighboring local plans, yielding local consistency without directly specifying the global structure of the full composition. As a result, locally compatible plans may still form an implausible route, task sequence, or temporal evolution. Existing methods improve global coherence by repeatedly propagating local consistency signals or by adding inference-time optimization, but these procedures become expensive as the number or dimensionality of local plans increases. We propose Coarse-to-Fine Compositional Diffusion (CoFi), an inference-time sampler that separates global structure formation from local detail refinement. CoFi first aligns local denoised estimates around a shared coarse structure, producing a global scaffold that captures the long-range task-level arrangement. It then diffuses this scaffold to an intermediate noise level and denoises it with the same pretrained local prior, restoring local fine structure while preserving the scaffold-induced global coherence. Across long-horizon robotic planning, panoramic image generation, and long video generation, CoFi not only improves both global coherence and local sample quality over prior compositional baselines, but also requires 2-8x fewer denoiser evaluations.

扩散模型长序列生成机器人规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。