用图结构融合视觉信息,实现无需训练的零样本导航。
T2Nav Algebraic Topology Aware Temporal Graph Memory and Loop Detection for ZeroShot Visual Navigation
- 构建含视觉信息的时序图记忆,动态更新环境拓扑。
- 在未知环境中实现稳定避障与精准回环检测,路径规划高效。
- 适合处理视觉相似但位置不同的目标实例导航任务。
在真实世界部署自主智能体面临挑战,尤其在导航方面,系统需应对未遇场景。传统学习方法依赖大量数据、频繁调参甚至重新训练,难以扩展且灵活性差。近期基础模型如大语言模型和视觉语言模型使系统能在不额外训练的情况下尝试新导航任务。然而,多数方法仅支持特定输入类型,推理能力有限,未能充分利用观察细节或空间结构。本文提出T2Nav,一种融合异构数据的零样本导航系统,采用基于图的推理机制。通过将视觉信息直接嵌入图结构并匹配环境,系统实现了探索与目标达成的良好平衡,具备鲁棒的障碍物规避、可靠的回环检测和高效的路径规划能力,同时避免重复探索模式。实验表明,该方法能有效适应未知环境,支持以目标物体实例参考图像为指令的导航,推动了实用化的零样本实例图像导航能力发展。
原文摘要 · Abstract (English)
Deploying autonomous agents in real world environments is challenging, particularly for navigation, where systems must adapt to situations they have not encountered before. Traditional learning approaches require substantial amounts of data, constant tuning, and, sometimes, starting over for each new task. That makes them hard to scale and not very flexible. Recent breakthroughs in foundation models, such as large language models and vision language models, enable systems to attempt new navigation tasks without requiring additional training. However, many of these methods only work with specific input types, employ relatively basic reasoning, and fail to fully exploit the details they observe or the structure of the spaces. Here, we introduce T2Nav, a zeroshot navigation system that integrates heterogeneous data and employs graph-based reasoning. By directly incorporating visual information into the graph and matching it to the environment, our approach enables the system to strike a good balance between exploration and goal attainment. This strategy allows robust obstacle avoidance, reliable loop closure detection, and efficient path planning while eliminating redundant exploration patterns. The system demonstrates flexibility by handling goals specified using reference images of target object instances, making it particularly suitable for scenarios in which agents must navigate to visually similar yet spatially distinct instances. Experiments demonstrate that our approach is efficient and adapts well to unknown environments, moving toward practical zero-shot instance-image navigation capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。