用多目标优化提升强化学习测试的多样性与效率
The Pursuit of Diversity: Multi-Objective Testing of Deep Reinforcement Learning Agents
- 联合优化失败概率与场景多样性,基于进化算法搜索
- 相比旧方法,失败类型多出40%~83%,且更快触发故障
- 适合安全关键领域测试,如自动驾驶与机器人
在安全关键领域测试深度强化学习(DRL)智能体需发现多样化的失败场景。现有工具INDAGO依赖单目标优化,仅追求最大失败数量,无法保证场景多样性或揭示不同错误类型。本文提出INDAGO-Nexus,一种多目标搜索方法,通过多目标进化算法结合多种多样性度量和帕累托前沿选择策略,同时优化失败可能性与测试场景多样性。我们在三类DRL智能体上评估:人形行走者、自动驾驶汽车(SDC)和停车代理。结果显示,平均而言,INDAGO-Nexus在SDC和停车场景中分别发现最多83%和40%更多唯一故障(测试有效性),且在所有智能体上将首次故障时间缩短达67%。
原文摘要 · Abstract (English)
Testing deep reinforcement learning (DRL) agents in safety-critical domains requires discovering diverse failure scenarios. Existing tools such as INDAGO rely on single-objective optimization focused solely on maximizing failure counts, but this does not ensure discovered scenarios are diverse or reveal distinct error types. We introduce INDAGO-Nexus, a multi-objective search approach that jointly optimizes for failure likelihood and test scenario diversity using multi-objective evolutionary algorithms with multiple diversity metrics and Pareto front selection strategies. We evaluated INDAGO-Nexus on three DRL agents: humanoid walker, self-driving car, and parking agent. On average, INDAGO-Nexus discovers up to 83% and 40% more unique failures (test effectiveness) than INDAGO in the SDC and Parking scenarios, respectively, while reducing time-to-failure by up to 67% across all agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。