自动生成高质量导航指令,提升智能体语言理解能力。
NavComposer: Composing Language Instructions for Navigation Trajectories through Action-Scene-Object Modularization
- 通过动作-场景-物体模块化分解与重组生成自然语言指令。
- 在多个轨迹上生成指令,准确率显著优于现有方法。
- 无需专家标注,适合大规模、跨场景的导航研究。
语言引导导航是具身人工智能的核心任务,使智能体能够理解语言指令并在复杂环境中导航。然而,专家提供的指令数量有限,而合成标注常质量不足,难以支撑大规模研究。为此,我们提出NavComposer,一个自动生成高质量导航指令的新框架。该框架显式分解动作、场景和物体等语义单元,并重新组合成自然语言指令。其模块化架构支持灵活集成先进方法,显式的语义单元提升了指令的丰富性与准确性。此外,该方法具备数据无关性,无需领域特定训练即可适配多样导航轨迹。为配合NavComposer,我们引入NavInstrCritic,一种无标注的评估系统,从对比匹配、语义一致性和语言多样性三个维度综合评价指令质量,克服了依赖人工标注的传统度量局限。通过解耦指令生成与评估过程,本方法实现更可扩展、通用的研究范式。大量实验提供了直接且实用的有效性证据。
原文摘要 · Abstract (English)
Language-guided navigation is a cornerstone of embodied AI, enabling agents to interpret language instructions and navigate complex environments. However, expert-provided instructions are limited in quantity, while synthesized annotations often lack quality, making them insufficient for large-scale research. To address this, we propose NavComposer, a novel framework for automatically generating high-quality navigation instructions. NavComposer explicitly decomposes semantic entities such as actions, scenes, and objects, and recomposes them into natural language instructions. Its modular architecture allows flexible integration of state-of-the-art techniques, while the explicit use of semantic entities enhances both the richness and accuracy of instructions. Moreover, it operates in a data-agnostic manner, supporting adaptation to diverse navigation trajectories without domain-specific training. Complementing NavComposer, we introduce NavInstrCritic, a comprehensive annotation-free evaluation system that assesses navigation instructions on three dimensions: contrastive matching, semantic consistency, and linguistic diversity. NavInstrCritic provides a holistic evaluation of instruction quality, addressing limitations of traditional metrics that rely heavily on expert annotations. By decoupling instruction generation and evaluation from specific navigation agents, our method enables more scalable and generalizable research. Extensive experiments provide direct and practical evidence for the effectiveness of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。