让机器人在陌生环境主动纠错,提升导航可靠性
Active Test-time Vision-Language Navigation
- 通过混合熵优化动态调节置信度,区分成功与失败路径
- 在REVERIE/R2R等数据集上,测试时性能超越基线10%以上
- 支持无人干预下的自反馈学习,适合真实场景部署
基于离线数据训练的视觉语言导航(VLN)策略在测试时面对陌生环境常表现下降,且缺乏外部交互与反馈。熵最小化虽能降低预测不确定性,但易因错误累积导致过度自信。为此,我们提出ATENA(主动测试时导航代理),一种通过周期性反馈实现人机协同的主动学习框架。ATENA学习在成功轨迹中增强置信度,在失败轨迹中降低置信度,改善不确定性校准。核心是混合熵优化:结合动作分布与伪专家分布(假设所选动作最优)计算熵,同时调控预测置信度与动作偏好。此外,提出自主动学习策略,使代理基于高置信预测评估自身导航结果。实验在挑战性基准REVERIE、R2R和R2R-CE上验证,ATENA有效缓解测试时分布偏移,在多种设置下均显著优于对比方法。
原文摘要 · Abstract (English)
Vision-Language Navigation (VLN) policies trained on offline datasets often exhibit degraded task performance when deployed in unfamiliar navigation environments at test time, where agents are typically evaluated without access to external interaction or feedback. Entropy minimization has emerged as a practical solution for reducing prediction uncertainty at test time; however, it can suffer from accumulated errors, as agents may become overconfident in incorrect actions without sufficient contextual grounding. To tackle these challenges, we introduce ATENA (Active TEst-time Navigation Agent), a test-time active learning framework that enables a practical human-robot interaction via episodic feedback on uncertain navigation outcomes. In particular, ATENA learns to increase certainty in successful episodes and decrease it in failed ones, improving uncertainty calibration. Here, we propose mixture entropy optimization, where entropy is obtained from a combination of the action and pseudo-expert distributions-a hypothetical action distribution assuming the agent's selected action to be optimal-controlling both prediction confidence and action preference. In addition, we propose a self-active learning strategy that enables an agent to evaluate its navigation outcomes based on confident predictions. As a result, the agent stays actively engaged throughout all iterations, leading to well-grounded and adaptive decision-making. Extensive evaluations on challenging VLN benchmarks-REVERIE, R2R, and R2R-CE-demonstrate that ATENA successfully overcomes distributional shifts at test time, outperforming the compared baseline methods across various settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。