arXiv:2606.26839cs.ROcs.CV2026-06被引 1

通过排序神经坍缩提升视觉导航的鲁棒性

Ordinal Neural Collapse as a Representation Prior for Visual Navigation

论文配图:Ordinal Neural Collapse as a Representation Prior for Visual Navigation
图 1 · 摘自论文原文
  • 利用导航动作的序关系构建有序表征空间
  • 在复杂路口场景下成功率提升18.7%,目标推进距离增加23%
  • 适合需要高鲁棒性的真实世界机器人导航任务

从视觉观测直接学习鲁棒的导航策略仍是基于视觉的机器人导航中的根本挑战。在端到端模仿学习中,视觉编码器与动作解码器通过单一动作损失联合优化,这对编码器仅提供间接监督信号,常导致编码器学习模糊、与动作无关的表征。这一问题在不同环境间显著的场景结构和外观差异,以及现实导航中普遍存在的视觉干扰下更加严重。此类与动作无关的特征使导航策略在模糊决策点产生不一致动作,导致导航失败。为此,我们提出 ORION(用于视觉导航的排序神经坍缩),显式根据导航动作的序结构组织编码器的表征空间。在目标导向导航中,从左远到右远的自身中心控制类别具有自然的序关系:相邻类别共享相似视觉上下文,语义对立类别则外观差异显著。我们促使类别表征沿单一判别轴顺序排列,同时抑制每类内部的轴外方差。预训练编码器被整合进基于扩散模型的导航框架,并进行端到端微调。在仿真与真实世界设置中的大量实验表明,ORION 在导航成功率和目标推进距离上均持续优于端到端及神经坍缩基线,在复杂多路交叉口等视觉挑战场景中表现尤为突出。

原文摘要 · Abstract (English)

Learning robust navigation policies directly from visual observations remains a fundamental challenge in vision-based robotic navigation. In end-to-end imitation learning approaches, the visual encoder and action decoder are jointly optimized using a single action loss, which provides only an indirect supervisory signal to the encoder. This indirect supervision frequently results in the encoder learning ambiguous, action-agnostic representations. The problem is further complicated by substantial variations in scene structure and appearance across diverse environments, as well as the prevalence of visual distractors inherent to real-world navigation settings. Such action-agnostic features cause the navigation policy to produce inconsistent actions at ambiguous decision points, leading to navigation failure. To overcome these limitations, we propose ORION (Ordinal Neural Collapse for Visual Navigation), a method that explicitly organizes the encoder's representation space according to the ordinal structure of navigation actions. In the context of goal-directed navigation, ego-centric control categories from Far Left to Far Right exhibit a natural ordinal relationship in which neighboring classes share similar visual contexts, while semantically opposing classes differ substantially in appearance. We encourage class representations to be arranged sequentially along a single discriminative axis, while suppressing off-axis variance within each class. The pretrained encoder is then integrated into a diffusion-based navigation framework, and the full pipeline is fine-tuned end-to-end. Extensive experiments in both simulation and real-world settings show that ORION consistently outperforms end-to-end and neural collapse baselines in navigation success rate and goal progress, with notable gains in visually challenging scenarios such as complex multi-way intersections.

视觉导航神经坍缩扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。