arXiv:2510.21809cs.CVcs.RO2025-10ICCV被引 3

让机器人导航同时学会用语言描述动作,提升可解释性。

Embodied Navigation with Auxiliary Task of Action Description Prediction

  • 在强化学习中加入动作描述作为辅助任务,增强决策透明度。
  • 在多模态导航任务中实现高精度导航与准确动作描述双达标。
  • 适合关注可解释AI与具身智能的科研人员与工程师。

近年来,室内环境下的多模态机器人导航受到广泛关注。然而,随着任务与方法日益复杂,动作决策系统趋于黑箱化,难以解释。可靠系统需具备解释自身决策的能力,但现有可解释系统性能常不及非可解释系统。本文提出将语言描述动作作为强化学习中的辅助任务,通过预训练视觉-语言模型的知识蒸馏解决缺乏真实标注数据的难题。我们在多种导航任务中全面评估该方法,证明其能在保持高导航性能的同时生成准确的动作描述,并在极具挑战性的语义音视频导航任务中达到当前最优表现。

原文摘要 · Abstract (English)

The field of multimodal robot navigation in indoor environments has garnered significant attention in recent years. However, as tasks and methods become more advanced, the action decision systems tend to become more complex and operate as black-boxes. For a reliable system, the ability to explain or describe its decisions is crucial; however, there tends to be a trade-off in that explainable systems can not outperform non-explainable systems in terms of performance. In this paper, we propose incorporating the task of describing actions in language into the reinforcement learning of navigation as an auxiliary task. Existing studies have found it difficult to incorporate describing actions into reinforcement learning due to the absence of ground-truth data. We address this issue by leveraging knowledge distillation from pre-trained description generation models, such as vision-language models. We comprehensively evaluate our approach across various navigation tasks, demonstrating that it can describe actions while attaining high navigation performance. Furthermore, it achieves state-of-the-art performance in the particularly challenging multimodal navigation task of semantic audio-visual navigation.

具身智能可解释性多模态导航

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。