arXiv:2510.11268cs.CV2025-10NeurIPS

提出类别向量,实现对图像分类器的高效灵活编辑。

Exploring and Leveraging Class Vectors for Classifier Editing

  • 用类别向量分离每个类别的表征调整,实现细粒度控制。
  • 通过类别向量算术可实现遗忘、自适应等高级编辑操作。
  • 无需重训,适合医疗影像与工业质检等场景快速迭代。

图像分类器在医学影像疾病检测和制造过程异常识别中至关重要。然而,经过大量训练后其行为固化,导致事后模型编辑困难,尤其在需遗忘特定类别或适应分布变化时。现有方法要么仅修正错误,要么代价高昂。为此,我们提出类别向量(Class Vectors),捕捉微调期间每个类别的特定表征调整。与任务向量编码权重空间的任务级变化不同,类别向量在隐空间解耦每个类别的适应性。我们证明类别向量能捕获类别的语义偏移,可通过沿这些向量引导隐特征或将其映射至权重空间来更新决策边界,实现分类器编辑。同时,类别向量的内在线性和正交性支持通过简单类别运算实现高效、灵活、高层的概念编辑。最后,我们在遗忘、环境适应、对抗防御及对抗触发优化等应用中验证了其有效性。

原文摘要 · Abstract (English)

Image classifiers play a critical role in detecting diseases in medical imaging and identifying anomalies in manufacturing processes. However, their predefined behaviors after extensive training make post hoc model editing difficult, especially when it comes to forgetting specific classes or adapting to distribution shifts. Existing classifier editing methods either focus narrowly on correcting errors or incur extensive retraining costs, creating a bottleneck for flexible editing. Moreover, such editing has seen limited investigation in image classification. To overcome these challenges, we introduce Class Vectors, which capture class-specific representation adjustments during fine-tuning. Whereas task vectors encode task-level changes in weight space, Class Vectors disentangle each class's adaptation in the latent space. We show that Class Vectors capture each class's semantic shift and that classifier editing can be achieved either by steering latent features along these vectors or by mapping them into weight space to update the decision boundaries. We also demonstrate that the inherent linearity and orthogonality of Class Vectors support efficient, flexible, and high-level concept editing via simple class arithmetic. Finally, we validate their utility in applications such as unlearning, environmental adaptation, adversarial defense, and adversarial trigger optimization.

分类器编辑类别向量无重训概念编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。