用扩散模型引导搜索,高效发现自动驾驶系统罕见故障。
Diffusion-Guided Search via Exponential Tilting (DiffTilt): An Application to Falsification of Safety-Critical Systems

- 通过指数倾斜扩散模型联合分布,聚焦故障相关行为
- 失败概率提升显著,优于传统条件采样方法
- 适合复杂系统验证,尤其在非STL规范下优势明显
在自主与网络物理系统中发现罕见安全关键故障是验证与测试的核心挑战。现有伪造方法依赖于条件采样策略,将环境与系统执行的联合分布分解,因而面临乘性稀疏效应:故障输入与故障轨迹同时稀少,导致穷举搜索代价高昂。本文提出DiffTilt,一种基于扩散模型联合分布的指数倾斜分布框架。我们证明扩散引导采样可精确解释为联合空间中的重要性采样,引导得分实现KL最优的概率质量重分配,集中于故障相关行为。进一步证明倾斜可严格放大失败概率,优于受乘性稀疏限制的条件采样。在此框架中,联合生成模型作为场景的可复用先验,无需真实反映待测系统;昂贵系统仿真仅用于学习评估函数以刻画场景质量,实现选择性与自适应使用。我们在ARCH-COMP基准上评估DiffTilt,提出新的拖拉机-挂车基准,展示当场景生成由明确规范而非奖励驱动时各方法的表现。所提方法相比最先进方法达到相当或更优的伪造性能,尤其在非STL公式定义规范时提升更显著。
原文摘要 · Abstract (English)
Discovering rare safety-critical failures in autonomous and cyber-physical systems is a fundamental challenge in verification and validation. Existing falsification approaches rely on conditional sampling strategies that factor the joint distribution over environments and system executions, and therefore suffer from multiplicative rarity effects: the simultaneous scarcity of failure-inducing inputs and failure-inducing traces makes exhaustive search prohibitively expensive. This paper develops DiffTilt, a distributional framework that exponentially tilts a diffusion model-induced joint distribution over environments and executions. We show that diffusion-guided sampling admits an exact interpretation as importance sampling in the joint space, where guidance scores induce a KL-optimal reallocation of probability mass towards failure-relevant behaviors. We further show that tilting provably amplifies failure probability and strictly outperforms conditional sampling, which is limited by multiplicative rarity. In this framework, the joint generative model serves as a reusable prior over scenarios and need not faithfully represent the system under test. Expensive system simulations are instead limited to learning a scoring function that characterizes scenario quality, enabling their selective and adaptive use. We study DiffTilt on ARCH-COMP benchmarks, and we propose an additional tractor-trailer benchmark showing the behavior of several approaches when scenario generation is guided by a well-defined specification rather than a reward. The proposed method achieves competitive or improved falsification performance compared to state-of-the-art approaches, with larger gains when specification definition is not limited to STL formulas.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。