arXiv:2606.20479cs.RO2026-06

通过轨迹一致性不确定性预测视觉语言导航失败,提升系统可靠性。

GroundControl: Anticipating Navigation Failures in Vision-Language Agents via Trajectory-Consistent Uncertainty Estimates

论文配图:GroundControl: Anticipating Navigation Failures in Vision-Language Agents via Trajectory-Consistent Uncertainty Estimates
图 1 · 摘自论文原文
  • 基于目标距离动态偏差构建轨迹一致性不确定度量,捕捉导航行为几何与时间不一致。
  • 在5个EB-Navigation数据集上,对GPT-4o模型的E-AURC达0.0024,接近最优排序。
  • 适用于需高可靠性导航的智能体部署,如机器人、自动驾驶等场景。

视觉语言导航代理在基准任务中表现良好,但常因可预测的轨迹级失效(如振荡、停滞或低效绕行)导致失败。可靠部署需依赖执行过程中能预判故障动态的不确定性信号,而非仅反映瞬时动作熵。我们提出「GroundControl」,一种基于轨迹一致性的不确定性估计器,定义为整个回合内名义目标导向距离-目标动态的统计偏离。该方法使用恒定速度卡尔曼滤波建模距离演化,结合归一化创新统计量与捕捉进展、单调性、路径效率和振荡行为的互补轨迹特征,生成反映导航行为几何与时间不一致性的不确定性得分,而非局部预测离散度。为独立评估不确定性质量,我们提出「选择性风险-覆盖导航(SRCN)」协议,利用风险-覆盖曲线及AURC/E-AURC指标衡量不确定性得分对失败或低效事件的排序能力。在五个EB-Navigation划分(共300个回合)中,轨迹一致性不确定性在成功导向的选择性风险下达到近似最优排序,GPT-4o模型加权平均E-AURC_SR为0.0024,显著优于熵、置信区间与启发式基线。在基于SPL的选择性评估中,GroundControl在所有模型与导航划分上均取得最低的AURC与E-AURC。结果表明,建模偏离目标导向动态的行为,可提供可解释且鲁棒的导航失败预判信号。

原文摘要 · Abstract (English)

Vision-language navigation agents achieve competitive average success on benchmark tasks, yet failures often arise through predictable trajectory-level breakdowns such as oscillation, stagnation, or inefficient detours. Reliable deployment, therefore, requires uncertainty signals that anticipate emerging failure dynamics during execution rather than reflect only instantaneous action entropy. We introduce \emph{GroundControl}, a trajectory-consistent uncertainty estimator defined as statistical deviation from nominal goal-directed distance-to-goal dynamics aggregated over an episode. GroundControl models distance evolution using a constant-velocity Kalman filter and combines normalized innovation statistics with complementary trajectory features capturing progress, monotonicity, path efficiency, and oscillatory behavior. The resulting uncertainty score reflects geometric and temporal inconsistency in navigation behavior rather than local prediction dispersion. To evaluate uncertainty quality independently of task success, we formalize \emph{Selective Risk--Coverage Navigation (SRCN)}, a protocol that measures how effectively an uncertainty score ranks episodes by failure or inefficiency using risk--coverage curves and AURC / E-AURC summaries. Across five EB-Navigation splits ($N=300$ episodes), trajectory-consistent uncertainty achieves near-oracle ordering under success-based selective risk, with weighted-average $\mathrm{E\text{-}AURC}_{\mathrm{SR}}=0.0024$ for the GPT-4o model, substantially outperforming entropy-, conformal-, and heuristic baselines. Under SPL-based selective evaluation, GroundControl consistently achieves the lowest AURC and E-AURC across models and navigation splits. These results show that modeling deviation from goal-directed dynamics provides an interpretable and robust signal for anticipating navigation failures in vision-language agents.

导航不确定性视觉语言强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。