从网络视频学习城市导航,揭示罕见但关键场景的长尾问题。
Exposing the Long-tail in Embodied Urban Navigation via Scalable Learning from In-the-Wild Videos

- 自动标注网络视频,生成带度量轨迹的导航语义数据。
- 模型在稀有场景下表现显著下降,暴露真实世界导航的长尾结构。
- 适合研究真实环境下的鲁棒性导航与数据偏差问题。
从真实世界数据学习具身城市导航策略受限于任务特定数据收集成本以及罕见但安全关键场景覆盖不足。为此,我们提出一种可扩展框架,通过网络规模的野外第一人称视频学习点目标城市导航,并系统揭示其长尾特征。该框架自动为未经筛选的网络视频标注度量轨迹和结构化导航语义,用于训练可解释的视觉-语言-动作策略。基于模型性能和感知-运动模式分布,我们刻画长尾特性,并采用反思分析诊断重复失败模式。在网页视频数据和真实城市导航任务上的实验表明,从非约束视频中有效迁移知识,揭示了超越整体导航性能的连贯长尾结构。
原文摘要 · Abstract (English)
Learning embodied urban navigation policies from real-world data is constrained by the cost of task-specific data collection and the limited coverage of rare yet safety-critical scenarios. To address these challenges, we present a scalable framework for learning point-goal urban navigation from web-scale in-the-wild egocentric videos while systematically exposing its long tail. The framework automatically annotates uncurated web videos with metric trajectories and structured navigation semantics, which are then used to train a vision-language-action policy for interpretable navigation planning. We characterize the long tail based on model performance and the distribution of perception-motion patterns, and employ reflection-based analysis to diagnose recurring failure modes. Experiments on web-video data and real-world urban navigation tasks demonstrate effective knowledge transfer from unconstrained videos and reveal coherent long-tail structures beyond aggregate navigation performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。