用连续流场替代传统路径预测,实现语言导航的高效闭环控制
CoFL: Continuous Flow Fields for Language-Conditioned Navigation
- 将导航任务转为在鸟瞰图上学习任意位置的运动向量
- 在50万对数据上训练,实测导航精度与安全性显著优于现有方法
- 零样本部署于真实场景,支持实时推理和闭环恢复
现有语言导航系统多依赖模块化流程或轨迹生成器,后者仅以单次起始条件监督轨迹。为此,我们提出CoFL,一种端到端策略,将鸟瞰图观测与语言指令映射为连续流场进行导航。CoFL将导航重构为基于工作空间的场学习,而非起始条件的轨迹预测:它在任意鸟瞰图位置学习局部运动向量,使每对场景-指令标注成为密集的空间控制监督信号。轨迹可通过数值积分任意起点生成,支持实时推演与闭环恢复。为支持大规模训练与评估,我们构建了超过50万组鸟瞰图-指令对数据集,每对由Matterport3D和ScanNet的语义地图生成流场与轨迹。在严格未见场景上评估,CoFL在导航精度与安全性上显著优于基于视觉-语言模型的规划器与轨迹生成策略,同时保持实时推理。最后,我们在多种布局的真实世界实验中零样本部署CoFL,实现了可行的闭环控制与高成功率。
原文摘要 · Abstract (English)
Existing language-conditioned navigation systems typically rely on modular pipelines or trajectory generators, but the latter use each scene--instruction annotation mainly to supervise one start-conditioned rollout. To address these limitations, we present CoFL, an end-to-end policy that maps a bird's-eye view (BEV) observation and a language instruction to a continuous flow field for navigation. CoFL reformulates navigation as workspace-conditioned field learning rather than start-conditioned trajectory prediction: it learns local motion vectors at arbitrary BEV locations, turning each scene--instruction annotation into dense spatial control supervision. Trajectories are generated from any start by numerical integration of the predicted field, enabling simple real-time rollout and closed-loop recovery. To enable large-scale training and evaluation, we build a dataset of over 500k BEV image--instruction pairs, each procedurally annotated with a flow field and a trajectory derived from semantic maps built on Matterport3D and ScanNet. Evaluating on strictly unseen scenes, CoFL significantly outperforms modular Vision-Language Model (VLM)-based planners and trajectory generation policies in both navigation precision and safety, while maintaining real-time inference. Finally, we deploy CoFL zero-shot in real-world experiments with BEV observations across multiple layouts, maintaining feasible closed-loop control and a high success rate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。