arXiv:2607.02222cs.ROcs.AI2026-07

用语言控制的局部流场实现连续导航,让机器人更自然地听懂指令走动。

CoFL-S: Spatially Queryable Sector Flow Fields for Local Language-Conditioned Navigation

论文配图:CoFL-S: Spatially Queryable Sector Flow Fields for Local Language-Conditioned Navigation
图 1 · 摘自论文原文
  • 基于局部可见区域生成语言驱动的流动场,直接预测连续动作路径
  • 在连续时间基准下,不同规划频率均优于传统离散动作基线
  • 支持零样本真实世界部署,适合追求流畅导航的机器人应用

视觉-语言导航日益关注高层指令推理、记忆与全局地图构建,而底层动作表示仍被忽视。本文提出 CoFL-S,一种低层视觉-语言-动作框架,通过预测机器人局部可见扇区内的语言条件流场,并滚动展开该场生成连续轨迹。为训练此低层表示,将每个 VLN-CE 任务转换为帧级局部监督:对齐子指令与匹配的动作、轨迹及密集流场目标。评估时引入连续时间 Habitat 基准,隔离低层动作接口,所有方法通过统一速度指令控制器执行,实现分解无关的闭环对比,突破 VLN-CE 中固定离散前进/转向的限制。在相同编码器与训练设置下,CoFL-S 在连续时间基准中各规划频率上均优于动作词元与动作块基线;零样本真实世界闭环部署也进一步验证其在仿真外的优势。

原文摘要 · Abstract (English)

Vision-Language Navigation has increasingly emphasized high-level instruction reasoning, memory, global map construction, and instruction decomposition, while the low-level action representation remains comparatively underexplored. We propose CoFL-S, a low-level vision-language-action framework that predicts a language-conditioned flow field over the robot's local visible sector and generates continuous trajectories by rolling out the predicted field. To train this low-level representation, we convert each VLN-CE episode, originally a whole-episode instruction paired with an action sequence, into frame-level local supervision with aligned sub-instructions and matched action, trajectory, and dense flow-field targets. For evaluation, we introduce a continuous-time Habitat benchmark that isolates low-level action interfaces from instruction decomposition and executes all methods through a shared velocity-command controller, enabling decomposition-independent closed-loop comparison across different planner frequencies rather than fixed discrete forward-and-turn transitions in VLN-CE. Under matched encoders and training settings, CoFL-S consistently outperforms action-token and action-chunk baselines across planner frequencies in the continuous-time Habitat benchmark, and zero-shot real-world closed-loop deployment further shows its advantage over both baselines beyond simulation.

视觉导航语言控制连续轨迹机器人决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。