arXiv:2508.11514cs.LG2025-08被引 1

提出双空间协同框架,提升决策智能体测试场景的多样性与关键性。

DiCriTest: Testing Scenario Generation for Decision-Making Agents Considering Diversity and Criticality

  • 分层表征参数空间,动态切换局部扰动与全局探索模式。
  • 在5个智能体上实现关键场景生成率提升56.23%,多样性更优。
  • 适合自动驾驶、机器人等高安全要求场景的测试验证。

决策智能体在动态环境中的部署日益增多,安全验证需求上升。尽管关键测试场景生成已成为有前景的验证方法,但有效平衡多样性与关键性仍是主要挑战,尤其因高维场景空间中易陷入局部最优。为此,本文提出一种双空间引导的测试框架,协调场景参数空间与智能体行为空间,以生成兼顾多样性和关键性的测试场景。在参数空间中,采用分层表征框架结合降维与多维子空间评估,高效定位多样化且关键的子空间,并动态协调局部扰动与全局探索两种生成模式,优化关键场景数量与多样性。在行为空间中,利用智能体-环境交互数据量化行为的关键性与多样性,自适应支持生成模式切换,形成闭环反馈,持续增强参数空间中的场景表征与探索能力。实验表明,该框架在五个决策智能体上平均提升关键场景生成率56.23%,并在新型参数-行为联合驱动指标下展现更强多样性,优于当前最先进基线方法。

原文摘要 · Abstract (English)

The growing deployment of decision-making agents in dynamic environments increases the demand for safety verification. While critical testing scenario generation has emerged as an appealing verification methodology, effectively balancing diversity and criticality remains a key challenge for existing methods, particularly due to local optima entrapment in high-dimensional scenario spaces. To address this limitation, we propose a dual-space guided testing framework that coordinates scenario parameter space and agent behavior space, aiming to generate testing scenarios considering diversity and criticality. Specifically, in the scenario parameter space, a hierarchical representation framework combines dimensionality reduction and multi-dimensional subspace evaluation to efficiently localize diverse and critical subspaces. This guides dynamic coordination between two generation modes: local perturbation and global exploration, optimizing critical scenario quantity and diversity. Complementarily, in the agent behavior space, agent-environment interaction data are leveraged to quantify behavioral criticality/diversity and adaptively support generation mode switching, forming a closed feedback loop that continuously enhances scenario characterization and exploration within the parameter space. Experiments show our framework improves critical scenario generation by an average of 56.23\% and demonstrates greater diversity under novel parameter-behavior co-driven metrics when tested on five decision-making agents, outperforming state-of-the-art baselines.

测试生成智能体多样性安全性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。