定位视觉语言导航中失败的具体能力短板,提升故障诊断可解释性。
Where Did It Go Wrong? Capability-Oriented Failure Attribution for Vision-and-Language Navigation Agents

- 通过种子选择与变异生成测试用例,动态覆盖复杂场景。
- 相比现有方法,发现更多失败案例并精准定位能力缺陷。
- 适合改进智能体可靠性,尤其在安全关键领域研究者使用。
在安全关键应用如视觉-语言导航(VLN)中,具身智能体依赖感知、记忆、规划、决策等多项相互关联的能力,导致故障难以定位和归因。现有测试方法多为系统级,无法揭示具体能力缺陷。本文提出一种面向能力的测试方法,结合(1)基于种子选择与变异的自适应测试用例生成,(2)能力判别器识别特定能力错误,(3)反馈机制将失败归因于具体能力并指导后续测试生成。实验表明,该方法比当前最优基线发现更多失败案例,并更准确地定位能力级缺陷,为改进具身智能体提供更具可解释性和可操作性的指导。
原文摘要 · Abstract (English)
Embodied agents in safety-critical applications such as Vision-Language Navigation (VLN) rely on multiple interdependent capabilities (e.g., perception, memory, planning, decision), making failures difficult to localize and attribute. Existing testing methods are largely system-level and provide limited insight into which capability deficiencies cause task failures. We propose a capability-oriented testing approach that enables failure detection and attribution by combining (1) adaptive test case generation via seed selection and mutation, (2) capability oracles for identifying capability-specific errors, and (3) a feedback mechanism that attributes failures to capabilities and guides further test generation. Experiments show that our method discovers more failure cases and more accurately pinpoints capability-level deficiencies than state-of-the-art baselines, providing more interpretable and actionable guidance for improving embodied agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。