让机器人通过对话构建持久导航状态,零样本提升复杂指令下的寻物能力。
SAIN: Structure-Aware Interactive Navigation with Active Dialogue Grounding for Mobile Robot

- 将对话内容转为结构化记忆:目标证据、路径记忆、物体候选标签
- 在VL-LN IIGN上成功率提至25.4%,路径长效率提至14.17
- 无需任务微调,适合真实场景下需主动提问的机器人导航
现有视觉语言导航多假设指令完整明确,但真实场景中人类指令常模糊或不完整,需机器人通过主动提问澄清。交互实例目标导航(IIGN)要求智能体在模糊类别指令下通过对话找到特定实例。现有方法仅将对话作为瞬时文本提示,未建立持续的空间或对象中心状态。本文提出SAIN,一种零样本框架,将主动对话转化为持久导航状态。它将对话答案整合为目标证据、路线级走廊记忆和物体候选标签,存入结构化值、房间、图与物体记忆,由统一策略用于前沿排序与目标逼近。在VL-LN IIGN基准上,相比最强基线,成功率从20.2提升至25.4,路径长效率从13.07升至14.17,且无需任务特定策略训练。结果验证了对话到状态转化在长程交互导航中的有效性。
原文摘要 · Abstract (English)
Most existing vision-language navigation tasks assume that instructions are complete and unambiguous. However, real-world robots often encounter natural human instructions that are ambiguous, underspecified, or incomplete, requiring them to resolve such uncertainties through active questioning. Interactive Instance Goal Navigation (IIGN) requires an embodied agent to find the specific instance under an ambiguous category-level instruction through active dialogue. However, existing dialogue-enabled methods often consume oracle answers as transient textual context for immediate decisions, rather than persistent spatial or object-centric structured state. We present SAIN, a zero-shot framework that turns active dialogue into persistent navigation state. Instead of consuming oracle answers as one-step text hints, SAIN compiles them into target evidence, route-level corridor memory, and object-candidate labels. These states are stored in structured value, room, graph, and object memories, then consumed by a unified policy for frontier ranking and final target approach. On the VL-LN IIGN benchmark, SAIN improves SR from 20.2 to 25.4 and SPL from 13.07 to 14.17 over the strongest reported dialogue-enabled baseline, while requiring no task-specific policy training. The results support dialogue-to-state conversion as an effective zero-shot mechanism for long-horizon interactive instance navigation. Project website: https://zorattc.github.io/SAIN/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。