用少步积分实现高精度视觉导航,解决生成式智能体延迟高的问题。
Rectified Schrödinger Bridge Matching for Few-Step Visual Navigation

- 利用熵正则参数统一建模随机与确定性路径,共享速度场结构
- 仅需3步积分即达94%方向相似度、92%成功率,远超传统方法
- 适合实时机器人控制场景,无需多阶段训练或知识蒸馏
视觉导航是具身智能的核心挑战,要求自主代理将高维感知信息转化为连续的长程动作轨迹。基于扩散模型和薛定谔桥(Schrödinger Bridge, SB)的生成策略虽能捕捉多模态动作分布,但因高方差随机传输需数十步积分,严重阻碍实时机器人控制。本文提出校正薛定谔桥匹配(Rectified Schrödinger Bridge Matching, RSBM),利用标准薛定谔桥(ε=1,最大熵传输)与确定性最优传输(ε→0,如条件流匹配)间共享的速度场结构,通过单一熵正则参数ε控制。证明两项关键结论:(1) 条件速度场函数形式在ε全谱下保持不变(速度结构不变性),支持单个网络适配所有正则强度;(2) 减小ε可线性降低条件速度方差,实现更稳定的粗粒度欧拉积分。依托学习到的条件先验缩短传输距离,RSBM在中间ε值下平衡多模态覆盖与路径直线性。实验表明,标准桥需≥10步收敛,而RSBM仅用3步即达到94%余弦相似度和92%成功率,无需知识蒸馏或多阶段训练,显著缩小高保真生成策略与具身智能低延迟需求之间的差距。
原文摘要 · Abstract (English)
Visual navigation is a core challenge in Embodied AI, requiring autonomous agents to translate high-dimensional sensory observations into continuous, long-horizon action trajectories. While generative policies based on diffusion models and Schrödinger Bridges (SB) effectively capture multimodal action distributions, they require dozens of integration steps due to high-variance stochastic transport, posing a critical barrier for real-time robotic control. We propose Rectified Schrödinger Bridge Matching (RSBM), a framework that exploits a shared velocity-field structure between standard Schrödinger Bridges ($\varepsilon=1$, maximum-entropy transport) and deterministic Optimal Transport ($\varepsilon\to 0$, as in Conditional Flow Matching), controlled by a single entropic regularization parameter $\varepsilon$. We prove two key results: (1) the conditional velocity field's functional form is invariant across the entire $\varepsilon$-spectrum (Velocity Structure Invariance), enabling a single network to serve all regularization strengths; and (2) reducing $\varepsilon$ linearly decreases the conditional velocity variance, enabling more stable coarse-step ODE integration. Anchored to a learned conditional prior that shortens transport distance, RSBM operates at an intermediate $\varepsilon$ that balances multimodal coverage and path straightness. Empirically, while standard bridges require $\geq 10$ steps to converge, RSBM achieves over 94% cosine similarity and 92% success rate in merely 3 integration steps -- without distillation or multi-stage training -- substantially narrowing the gap between high-fidelity generative policies and the low-latency demands of Embodied AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。