提出实时可控的生成式机器人策略,反应更快更灵活。
Guided Streaming Stochastic Interpolant Policy

- 基于随机插值模型推导最优引导项,实现动态目标追踪
- 实测在动态环境中反应速度显著优于传统分块方法
- 支持零样本适配与训练后优化,适合真实场景部署
推理时引导对于在不重新训练的情况下引导生成式机器人策略以应对动态目标至关重要,但现有方法多局限于分块架构,存在高延迟且缺乏测试时偏好对齐或避障所需的响应性。本文通过分析价值函数的时间演化(基于后向柯尔莫哥洛夫方程),形式化推导出随机插值(SI)的最优引导项,建立理论保证采样于目标分布的修正漂移项。将该框架应用于实时控制,提出流式随机插值策略(SSIP),其推广了确定性流式策略(SFP)。将此引导律与流式架构结合,实现快速且具反应性的控制。为满足多样化部署需求,提出两种互补机制:无需训练的在线梯度计算方法(STEG),实现零样本适应;以及基于训练的条件评判器引导(CCG),实现推理成本摊销。实证评估表明,所提引导流式方法在反应性上显著优于传统分块策略,并在动态非结构化环境中提供更优、物理合理的引导效果。
原文摘要 · Abstract (English)
Inference-time guidance is essential for steering generative robot policies toward dynamic objectives without retraining, yet existing methods are largely confined to chunk-based architectures that exhibit high latency and lack the reactivity needed for test-time preference alignment or obstacle avoidance. In this work, we formally derive the optimal guidance term for Stochastic Interpolants (SI) by analyzing the value function's time evolution via the Backward Kolmogorov Equation, establishing a modified drift that theoretically guarantees sampling from a target distribution. We apply this framework to real-time control through the Streaming Stochastic Interpolant Policy (SSIP), which generalizes the deterministic Streaming Flow Policy (SFP). Unifying this guidance law with the streaming architecture enables fast and reactive control. To support diverse deployment needs, we propose two complementary mechanisms: training-free Stochastic Trajectory Ensemble Guidance (STEG) that computes gradients on-the-fly for zero-shot adaptation, and training-based Conditional Critic Guidance (CCG) for amortized inference. Empirical evaluations demonstrate that our guided streaming approach significantly outperforms conventional chunk-based policies in reactivity and provides superior, physically valid guidance for dynamic, unstructured environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。