arXiv:2507.06747cs.ROcs.CV2025-07被引 10

让机器人在复杂环境中自主寻找任意物体,支持长距离导航与动态适应。

LOVON: Legged Open-Vocabulary Object Navigator

  • 结合大模型与开放词汇视觉检测,实现分层任务规划与目标识别。
  • 在真实机器人上完成多阶段搜索与导航,成功率达92%以上。
  • 适配多种轮足机器人,具备即插即用的部署优势。

开放世界中的物体导航对机器人系统仍是重大挑战,尤其在需要长时程任务执行时,需同时处理开放词汇物体检测与高层任务规划。传统方法难以有效整合这些模块,限制了其应对复杂、远距离导航任务的能力。本文提出LOVON框架,将大语言模型(LLMs)用于分层任务规划,并结合专为动态非结构化环境设计的开放词汇视觉检测模型,实现高效的长距离物体导航。针对实际场景中的视觉抖动、盲区和目标临时丢失等问题,设计了拉普拉斯方差滤波等专用解决方案。开发了功能性执行逻辑,确保机器人在自主导航、任务适应和鲁棒完成方面的能力。大量实验表明,该系统能成功完成涉及实时检测、搜索与导航至开放词汇动态目标的长序列任务。此外,在Unitree Go2、B2和H1-2等多种轮足机器人上的真实世界测试,验证了LOVON的兼容性与优异的即插即用特性。

原文摘要 · Abstract (English)

Object navigation in open-world environments remains a formidable and pervasive challenge for robotic systems, particularly when it comes to executing long-horizon tasks that require both open-world object detection and high-level task planning. Traditional methods often struggle to integrate these components effectively, and this limits their capability to deal with complex, long-range navigation missions. In this paper, we propose LOVON, a novel framework that integrates large language models (LLMs) for hierarchical task planning with open-vocabulary visual detection models, tailored for effective long-range object navigation in dynamic, unstructured environments. To tackle real-world challenges including visual jittering, blind zones, and temporary target loss, we design dedicated solutions such as Laplacian Variance Filtering for visual stabilization. We also develop a functional execution logic for the robot that guarantees LOVON's capabilities in autonomous navigation, task adaptation, and robust task completion. Extensive evaluations demonstrate the successful completion of long-sequence tasks involving real-time detection, search, and navigation toward open-vocabulary dynamic targets. Furthermore, real-world experiments across different legged robots (Unitree Go2, B2, and H1-2) showcase the compatibility and appealing plug-and-play feature of LOVON.

机器人导航开放词汇大模型轮足机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。