arXiv:2512.20083cs.SEcs.RO2025-12被引 1

用变异测试检测智能体任务规划中的低效决策,提升部署可靠性。

Detecting Non-Optimal Decisions of Embodied Agents via Diversity-Guided Metamorphic Testing

  • 通过多样性引导的变异测试,识别规划中非最优行为。
  • 平均检测率达31.9%,比最优基线高16.8%相对提升。
  • 适用于验证各类规划模型,尤其适合资源受限场景。

随着具身智能体向真实世界部署推进,确保决策最优性对资源受限应用至关重要。现有评估方法主要关注功能正确性,忽视了生成计划的非功能性最优性,可能导致性能下降与资源浪费。本文定义并形式化了非最优决策(NoDs)问题:智能体成功完成任务但效率低下。提出NoD-DGMT框架,通过多样性引导的变异测试系统检测具身智能体任务规划中的非最优行为。核心思想是:最优规划器在特定变换下应保持行为不变性。设计四种新型变异关系,涵盖位置绕行次优性、动作最优完整性、条件细化单调性与场景扰动不变性。引入多样性引导选择策略,主动选取覆盖不同违规类别的测试用例,避免冗余评估并保证全面多样性覆盖。在AI2-THOR模拟器上对四个先进规划模型的实验表明,NoD-DGMT平均检测率达31.9%,多样性引导过滤器使检测率提升4.3%,多样性得分提高3.3。显著优于六种基线方法,相对最佳基线提升16.8%,且在不同模型架构和任务复杂度下表现一致优越。

原文摘要 · Abstract (English)

As embodied agents advance toward real-world deployment, ensuring optimal decisions becomes critical for resource-constrained applications. Current evaluation methods focus primarily on functional correctness, overlooking the non-functional optimality of generated plans. This gap can lead to significant performance degradation and resource waste. We identify and formalize the problem of Non-optimal Decisions (NoDs), where agents complete tasks successfully but inefficiently. We present NoD-DGMT, a systematic framework for detecting NoDs in embodied agent task planning via diversity-guided metamorphic testing. Our key insight is that optimal planners should exhibit invariant behavioral properties under specific transformations. We design four novel metamorphic relations capturing fundamental optimality properties: position detour suboptimality, action optimality completeness, condition refinement monotonicity, and scene perturbation invariance. To maximize detection efficiency, we introduce a diversity-guided selection strategy that actively selects test cases exploring different violation categories, avoiding redundant evaluations while ensuring comprehensive diversity coverage. Extensive experiments on the AI2-THOR simulator with four state-of-the-art planning models demonstrate that NoD-DGMT achieves violation detection rates of 31.9% on average, with our diversity-guided filter improving rates by 4.3% and diversity scores by 3.3 on average. NoD-DGMT significantly outperforms six baseline methods, with 16.8% relative improvement over the best baseline, and demonstrates consistent superiority across different model architectures and task complexities.

智能体测试方法优化检测具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。