arXiv:2609.05596cs.CVcs.RO2026-09

让AI导航助手学会在对的时间说对的话,提升视障者出行安全。

Time-Aware Assistive Navigation

论文配图:Time-Aware Assistive Navigation
图 1 · 摘自论文原文
  • 通过预测指令背后的理由,改进模型的时机判断能力。
  • 在多种场景下,改进后模型显著提升响应及时性与准确性。
  • 适合研究具身智能、多模态大模型与辅助系统开发人员。

当前语言模型很少规划何时做出实时回应。在引导视障者穿越复杂城市环境时,准确且及时的支持至关重要,错误时机的反馈可能分散注意力或增加认知负担。为此,我们构建了一个大规模多模态基准,用于评估基于多模态大模型(MLLM)的主动导航代理在第一人称户外环境中的表现。实验发现,即使经过大量数据微调,现成的MLLM在提供安全且及时的导航指令方面仍存在根本缺陷。我们提出一种简单有效的改进方法:直接监督模型预测每条指令背后的理由,该方法在开环、闭环及模拟到真实环境迁移测试中均取得显著性能提升。然而,分析揭示了时间推理、安全关键物体感知以及关系和距离理解等方面的持续挑战。为推动可扩展辅助代理的发展,我们将公开仿真环境、基准数据集与代码(项目网站:https://timeli-icra.github.io/)。

原文摘要 · Abstract (English)

Can interactive vision-and-language agents learn not just what to say but also \textbf{\textit{when}} to say it? Current language models rarely plan over whether and when to realize a real-time response to a user. However, providing accurate and timely support for human decision-making, such as when guiding visually impaired individuals through urban environments, requires careful real-time responsiveness--poorly timed responses can distract users or add unnecessary cognitive load. As a machine intelligence challenge for Multimodal Large Language Model (MLLM)-based agents, we introduce a large-scale multimodal benchmark for an egocentric, assistive navigation task in complex outdoor environments. Using this benchmark, we uncover a fundamental limitation of off-the-shelf MLLMs in delivering safe and time-sensitive navigation instructions, even with model fine-tuning on substantial amounts of data. We then demonstrate that a simple yet effective modification of the model, including direct supervision to predict the underlying reason for each instruction, yields significant performance gains across open-loop, closed-loop, and sim-to-real generalization settings. However, our analysis highlights persistent challenges in temporal reasoning, safety-critical object awareness, and relational and distance understanding. To advance the development of scalable assistive agents, we will release our simulation, benchmark, and code (available at the project website: https://timeli-icra.github.io/).

导航助手多模态时间感知视障辅助

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。