arXiv:2604.25161cs.MAcs.AI2026-04ACL

定位视觉语言导航中失败的具体能力短板,提升故障诊断可解释性。

Where Did It Go Wrong? Capability-Oriented Failure Attribution for Vision-and-Language Navigation Agents

论文配图:Where Did It Go Wrong? Capability-Oriented Failure Attribution for Vision-and-Language Navigation Agents
图 1 · 摘自论文原文
  • 通过种子选择与变异生成测试用例,动态覆盖复杂场景。
  • 相比现有方法,发现更多失败案例并精准定位能力缺陷。
  • 适合改进智能体可靠性,尤其在安全关键领域研究者使用。

在安全关键应用如视觉-语言导航(VLN)中,具身智能体依赖感知、记忆、规划、决策等多项相互关联的能力,导致故障难以定位和归因。现有测试方法多为系统级,无法揭示具体能力缺陷。本文提出一种面向能力的测试方法,结合(1)基于种子选择与变异的自适应测试用例生成,(2)能力判别器识别特定能力错误,(3)反馈机制将失败归因于具体能力并指导后续测试生成。实验表明,该方法比当前最优基线发现更多失败案例,并更准确地定位能力级缺陷,为改进具身智能体提供更具可解释性和可操作性的指导。

原文摘要 · Abstract (English)

Embodied agents in safety-critical applications such as Vision-Language Navigation (VLN) rely on multiple interdependent capabilities (e.g., perception, memory, planning, decision), making failures difficult to localize and attribute. Existing testing methods are largely system-level and provide limited insight into which capability deficiencies cause task failures. We propose a capability-oriented testing approach that enables failure detection and attribution by combining (1) adaptive test case generation via seed selection and mutation, (2) capability oracles for identifying capability-specific errors, and (3) a feedback mechanism that attributes failures to capabilities and guides further test generation. Experiments show that our method discovers more failure cases and more accurately pinpoints capability-level deficiencies than state-of-the-art baselines, providing more interpretable and actionable guidance for improving embodied agents.

视觉语言导航故障归因智能体评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。