arXiv:2606.21216cs.RO2026-06

用预训练ViT的每块标量信息,实现真实世界快速导航

A scalar per patch from pre-trained ViTs enables fast moving navigation in the real world

论文配图:A scalar per patch from pre-trained ViTs enables fast moving navigation in the real world
图 1 · 摘自论文原文
  • 提取ViT每块的标量特征,高效压缩视觉信息
  • 仅用标量特征实现24公里真实场景导航,效果接近全模型
  • 适合关注视觉压缩与机器人导航融合的研究者

真实世界机器人导航依赖计算机视觉组件,通常为预训练视觉编码器。这些编码器对机器人性能至关重要,其能力不仅源于下游任务训练,更依赖于计算机视觉预训练任务中的辅助损失。本研究在真实建筑中完成966次静态目标导航实验,累计24公里,系统评估主流视觉编码器在实际条件下的表现。研究探索异构多教师蒸馏,构建具备多种互补能力的编码器;通过有原则且空间有意义的瓶颈机制,研究编码器所需信息量,发现该方法可催生与操作性相关的可解释特征。此外,证明仅基于RGB数据训练策略无法最优利用视觉特征,通过微调在特权信息上预训练的策略可显著提升性能。本研究全面揭示了计算机视觉在真实世界导航中的关键要素。

原文摘要 · Abstract (English)

Trained policies for real-world robotics rely on computer vision components, typically in the form of pre-trained visual encoders. These encoders are an essential component and it has been shown that their power does not emerge from training on robotics downstream losses alone. Pre-training with auxiliary losses in the form of computer-vision pre-text tasks is a defining factor and heavily conditions agent performance in robotics tasks. In this unprecedented large-scale study, we ran 966 navigation episodes of static point goal navigation in a real-world building for 24km and asked which components really matter for the computer vision aspects of robotics: we evaluate state-of-the art visual encoders in realistic conditions. We explore the usefulness of heterogeneous multi-teacher distillation leading to encoders with multiple different and complementary skills. We investigate how much information from these encoders is necessary for robotics by bottlenecking them in a principled and spatially useful way and we show that this leads to the emergence of interpretable features linked to affordances. We also argue that training policies on RGB data alone does not lead to an optimal usage of visual features and show this by finetuning policies pre-trained on privileged information. All in all, we paint a more complete picture of what aspects of computer vision are relevant for real-world navigation.

视觉编码器机器人导航特征压缩ViT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。