arXiv:2603.15656cs.LGcs.AI2026-03中稿 · CVPR被引 3

用少量干净样本修复模型对噪声数据的错误反应。

Attribution-Guided Model Rectification of Unreliable Neural Network Behaviors

  • 基于梯度归因定位错误根源层,实现精准修正。
  • 仅需一个干净样本即可完成有效修复,计算成本极低。
  • 适用于对抗攻击、虚假关联等不可靠行为修复。

神经网络在受损样本的非鲁棒特征上表现出不可靠行为,导致性能下降。由于模型的黑箱特性,修复此类问题通常需要大量数据清洗和重新训练,带来巨大计算与人工成本。本文提出一种基于秩一模型编辑的归因引导修复框架,能有效定位并修正模型的不可靠行为。我们区分了现有模型编辑与本工作设定,提出在不牺牲模型性能的前提下,减少对清洁样本依赖的修复方案。进一步揭示了不同层编辑能力异质性带来的瓶颈,提出一种归因引导的层定位方法,量化各层可编辑性,识别出最可能导致错误的核心层。大量实验表明,该方法在修复神经后门、虚假相关和特征泄露等问题上表现优异,仅需单个干净样本即可达成目标,具有很强的实用性。

原文摘要 · Abstract (English)

The performance of neural network models deteriorates due to their unreliable behavior on non-robust features of corrupted samples. Owing to their opaque nature, rectifying models to address this problem often necessitates arduous data cleaning and model retraining, resulting in huge computational and manual overhead. In this work, we leverage rank-one model editing to establish an attribution-guided model rectification framework that effectively locates and corrects model unreliable behaviors. We first distinguish our rectification setting from existing model editing, yielding a formulation that corrects unreliable behavior while preserving model performance and reducing reliance on large budgets of cleansed samples. We further reveal a bottleneck of model rectifying arising from heterogeneous editability across layers. To target the primary source of misbehavior, we introduce an attribution-guided layer localization method that quantifies layer-wise editability and identifies the layer most responsible for unreliabilities. Extensive experiments demonstrate the effectiveness of our method in correcting unreliabilities observed for neural Trojans, spurious correlations and feature leakage. Our method shows remarkable performance by achieving its editing objective with as few as a single cleansed sample, which makes it appealing for practice.

模型修复归因分析鲁棒性神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。