通过神经元激活状态自动推断神经网络的正式性质
Prophecy: Inferring Formal Properties from Neuron Activations
- 基于隐藏层神经元激活值或开关状态提取逻辑规则
- 可推导出预测结果属于特定类别的形式化条件
- 适用于大模型解释、运行时监控与错误修复
我们提出 Prophecy,一个用于自动推断前馈神经网络形式性质的工具。其核心思想是:前馈网络的大部分逻辑体现在内部层神经元的激活状态中。Prophecy 通过提取以神经元激活(数值或开/关状态)为前提条件的规则,来推断某些期望的输出属性,例如预测结果属于特定类别。这些规则揭示了隐藏层中捕获的、能预示期望输出行为的网络性质。本文介绍了该工具的架构,展示了其在多种模型和输出属性上的应用效果,并概述了其在生成形式化解释、组合验证、运行时监控、修复等任务中的潜力。此外,我们还展示了其在大型视觉-语言模型时代的新颖应用前景。
原文摘要 · Abstract (English)
We present Prophecy, a tool for automatically inferring formal properties of feed-forward neural networks. Prophecy is based on the observation that a significant part of the logic of feed-forward networks is captured in the activation status of the neurons at inner layers. Prophecy works by extracting rules based on neuron activations (values or on/off statuses) as preconditions that imply certain desirable output property, e.g., the prediction being a certain class. These rules represent network properties captured in the hidden layers that imply the desired output behavior. We present the architecture of the tool, highlight its features and demonstrate its usage on different types of models and output properties. We present an overview of its applications, such as inferring and proving formal explanations of neural networks, compositional verification, run-time monitoring, repair, and others. We also show novel results highlighting its potential in the era of large vision-language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。