arXiv:2511.07293cs.LOcs.AI2025-11

给神经网络的置信度做形式化验证,提升模型决策可信度。

Formal Reasoning About Confidence and Automated Verification of Neural Networks

  • 设计通用语法描述置信度相关规范
  • 通过添加层实现统一验证,支持主流工具
  • 在8870个基准上验证,最大模型138M参数

过去十年中,大量研究关注神经网络的鲁棒性,即输入轻微扰动时输出是否不变。但多数方法忽略了模型对输出的置信度。本文提出一个形式化框架,同时处理置信度与鲁棒性。设计了一种简洁而表达力强的语法,用于捕捉多种置信度相关规格。提出一种统一的新技术:通过向神经网络添加少量层,使任意先进验证工具都能统一处理所有语法实例。在包含8870个基准的大规模实验中进行评估,最大网络达13800万参数,结果显著优于传统编码方法。

原文摘要 · Abstract (English)

In the last decade, a large body of work has emerged on robustness of neural networks, i.e., checking if the decision remains unchanged when the input is slightly perturbed. However, most of these approaches ignore the confidence of a neural network on its output. In this work, we aim to develop a generalized framework for formally reasoning about the confidence along with robustness in neural networks. We propose a simple yet expressive grammar that captures various confidence-based specifications. We develop a novel and unified technique to verify all instances of the grammar in a homogeneous way, viz., by adding a few additional layers to the neural network, which enables the use any state-of-the-art neural network verification tool. We perform an extensive experimental evaluation over a large suite of 8870 benchmarks, where the largest network has 138M parameters, and show that this outperforms ad-hoc encoding approaches by a significant margin.

神经网络形式化验证置信度鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。