arXiv:2606.21014cs.RO2026-06

用费曼-卡茨采样让机器人轨迹生成更安全高效

BayesFP: Posterior Estimation for Flow-Based Policies via Feynman-Kac Sampling

论文配图:BayesFP: Posterior Estimation for Flow-Based Policies via Feynman-Kac Sampling
图 1 · 摘自论文原文
  • 将轨迹生成转为贝叶斯后验采样,利用专家数据作先验
  • 在不重训练的情况下,实现对扩散与流匹配模型的统一推理采样
  • 可在推理时动态处理非凸障碍物,适合零样本任务部署

机器人需生成符合学习到的专家行为、满足安全约束且在推理时才给出任务目标的轨迹。本文将预训练扩散模型和流匹配策略的约束轨迹生成问题建模为贝叶斯后验采样:以学习到的示范分布为先验,通过推理时由代价函数导出的似然项调整方向,使轨迹趋向可行且最优。为在不重训练基础策略的前提下采样该后验分布,我们引入原本用于扩散模型的费曼-卡茨校正框架,并将其扩展至确定性流匹配策略。由此得到一个统一的、仅在推理阶段使用、无需重训练的采样方法,适用于扩散与流匹配策略。我们在预训练的Diffusion Policy、GR00T-N1.6和π_{0.5}检查点上验证了该方法,在模拟与真实世界操作任务中均表现优异,包括在推理时引入非凸障碍物的情况,并在零样本任务上超越基线π_{0.5}。

原文摘要 · Abstract (English)

Robots must generate trajectories that remain faithful to learned expert behavior while satisfying safety constraints and task-specific objectives specified only at inference time. We formulate constrained trajectory generation for pretrained diffusion and flow-matching policies as Bayesian posterior sampling, with the learned demonstration distribution as a prior and an inference-time, cost-derived likelihood tilting it toward feasible, optimal trajectories. To sample from this posterior without any retraining of the base policy, we leverage the Feynman--Kac corrector framework, originally formulated for diffusion models, and extend it to deterministic flow-matching policies. The result is a unified, inference-time, retraining-free sampler for diffusion and flow policies. We validate the approach on pretrained Diffusion Policy, GR00T-N1.6, and $π_{0.5}$ checkpoints across simulated and real-world manipulation tasks, including planning around non-convex obstacles introduced at inference time, and show improvements over the base $π_{0.5}$ on zero-shot tasks.

机器人控制轨迹生成贝叶斯推断流匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。