用代理系统指导真实系统测试,高效发现罕见故障。
Coverage Aware Active Evaluation for Failure Discovery with Paired Systems

- 通过残差建模修正代理系统的失败信号,构建目标系统风险预测器。
- 在自动驾驶等任务中,故障发现数量是基线方法的2倍。
- 适合资源受限下需高效发现多样严重故障的场景。
自主系统可能以稀有且多样的方式失效,受限于测试预算,真实世界中的故障发现十分困难。尽管模拟器、低保真系统或相关策略等廉价代理系统可广泛采样以发现故障,但其故障常因仿真到现实及系统间差异而无法迁移至真实环境。核心挑战在于如何有效利用代理系统信息,准确预测目标系统的严重故障。本文提出一种自适应故障发现方法,结合代理评估与有限的目标系统结果,引导目标系统测试场景的选择。该方法通过受控制变量启发的残差建模,校正代理失败信号,学习局部目标风险预测器。为同时追求高概率与多样性,将预测器与支持感知互信息目标结合,偏好真实且充分支持的区域,同时扩展对故障模式的覆盖。在自动驾驶、操作与四足速度跟踪任务中,本方法发现的故障数量比随机采样和主动学习基线高出最多2倍,包含被其他方法遗漏的严重且多样的故障。
原文摘要 · Abstract (English)
Autonomous systems can fail in rare and heterogeneous ways, making real-world failure discovery difficult under limited testing budgets. Although cheaper proxies such as simulators, lower-fidelity systems, or related policies can be sampled extensively to find failures, proxy failures often do not transfer to the real world due to sim-to-real and system-to-system gaps. The key challenge is therefore to effectively leverage proxy system information for accurate prediction of severe target system failures. We propose an adaptive failure discovery method that combines proxy evaluations with limited target system results to guide scenario selection for target system testing. Our method learns a local predictor of target risk by correcting proxy failure signals using control-variate-inspired residual modeling. To find failures that are both likely and diverse, we combine this predictor with a support-aware mutual-information objective that favors realistic, well-supported regions while expanding coverage across failure modes. Across autonomous driving, manipulation, and quadruped velocity-tracking tasks, our method discovers up to 2$\times$ as many failures as random sampling and active-learning baselines, including severe and diverse failures missed by competing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。