用大模型自动发现决策系统的关键测试场景。
Exploring Critical Testing Scenarios for Decision-Making Policies: An LLM Approach
- 基于大模型生成-测试-反馈循环,利用提示工程激发推理能力。
- 多尺度生成策略提升测试效率,发现更多关键与多样场景。
- 在五个基准上优于基线,适合自动驾驶等高可靠性领域测试。
决策策略在自动驾驶、机器人等领域取得显著进展,但其可靠性仍受关键场景测试不足的威胁。现有方法面临测试效率低、场景多样性有限等问题,源于策略与环境的复杂性。本文提出一种可适配的大语言模型(LLM)驱动在线测试框架,设计‘生成-测试-反馈’流程,结合模板化提示工程,利用LLM的世界知识与推理能力。进一步提出多尺度场景生成策略,弥补LLM精细调整能力的不足,提升测试效率。在五个主流基准上的实验表明,该方法显著优于基线,在挖掘关键与多样化场景方面表现更优。结果表明,LLM驱动方法在推进决策策略测试方面具有巨大潜力。
原文摘要 · Abstract (English)
Recent advances in decision-making policies have led to significant progress in fields such as autonomous driving and robotics. However, testing these policies remains crucial with the existence of critical scenarios that may threaten their reliability. Despite ongoing research, challenges such as low testing efficiency and limited diversity persist due to the complexity of the decision-making policies and their environments. To address these challenges, this paper proposes an adaptable Large Language Model (LLM)-driven online testing framework to explore critical and diverse testing scenarios for decision-making policies. Specifically, we design a "generate-test-feedback" pipeline with templated prompt engineering to harness the world knowledge and reasoning abilities of LLMs. Additionally, a multi-scale scenario generation strategy is proposed to address the limitations of LLMs in making fine-grained adjustments, further enhancing testing efficiency. Finally, the proposed LLM-driven method is evaluated on five widely recognized benchmarks, and the experimental results demonstrate that our method significantly outperforms baseline methods in uncovering both critical and diverse scenarios. These findings suggest that LLM-driven methods hold significant promise for advancing the testing of decision-making policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。