arXiv:2602.20114cs.CVcs.AI2026-02中稿 · CoLLAs 2026

首个针对视觉Transformer的遗忘能力评测,揭示现有算法在模型上的表现差异。

Benchmarking Unlearning for Vision Transformers

  • 对比多种视觉Transformer架构与遗忘算法,设计统一评测流程
  • 发现利用训练数据记忆性可显著提升遗忘性能,优于传统方法
  • 适合关注AI安全、隐私保护及模型可解释性的研究者参考

机器遗忘(MU)是指在训练后移除错误、偏见或泄露敏感信息的训练样本影响的能力,现被视为构建安全公平AI的关键。尽管视觉变换器(VTs)已逐渐替代卷积神经网络(CNNs),但当前的遗忘研究仍集中于CNNs,缺乏针对VTs的评测基准。本文首次构建了面向不同家族(ViT、Swin-T、DINOv2)和容量的视觉变换器遗忘评测体系,涵盖多数据集(评估规模与复杂度影响)、多种遗忘算法(代表不同方法范式)以及单次与持续遗忘协议。特别关注基于训练数据记忆性的遗忘算法,因近期研究表明其可显著提升以往最优算法性能。研究还分析了视觉变换器的记忆特性,并评估不同记忆代理的影响。采用统一指标衡量遗忘质量与测试/保留数据上的准确率。该工作为现有及未来遗忘算法在视觉变换器上的比较提供了可复现、公平、全面的基准,首次揭示了主流算法在视觉变换器场景下的实际表现,确立了有前景的性能基线。

原文摘要 · Abstract (English)

Machine unlearning (MU) refers to the post-training capability to remove (the influence of) training examples that are incorrect, biased, or leak sensitive/private information. MU is now widely regarded as critical for building safe and fair AI. In parallel, research into transformer architectures for computer vision has been highly successful: Vision Transformers (VTs) increasingly emerge as strong alternatives to CNNs. Yet, MU research for vision tasks has largely centered on CNNs, not VTs. While MU benchmarks have been developed for LLMs, diffusion models, and CNNs, none currently exist for VTs. This work is the first to attempt this, benchmarking MU algorithm performance across different VT families (ViT, Swin-T, and DINOv2) and at different capacities. The work employs (i) different datasets, selected to assess the impacts of dataset scale and complexity; (ii) different MU algorithms, selected to represent fundamentally different approaches for MU; and (iii) both single-shot and continual unlearning protocols. Additionally, it focuses on benchmarking MU algorithms that leverage training data memorization, since leveraging memorization has been recently discovered to significantly improve the performance of previously SOTA algorithms. En route, the work characterizes how VTs memorize training data relative to CNNs, and assesses the impact of different memorization proxies on performance. The benchmark uses unified evaluation metrics that capture two complementary notions of forget quality along with accuracy on unseen (test) data and on retained data. Overall, this work offers a benchmarking basis, enabling reproducible, fair, and comprehensive comparisons of existing (and future) MU algorithms on VTs. Importantly, for the first time, it sheds light on how well existing algorithms work in VT settings, establishing a promising reference performance baseline.

机器遗忘视觉Transformer模型安全评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。