arXiv:2512.12469cs.LG2025-12

用极少标注让神经表示中的特定概念可解释可控。

Sparse Concept Anchoring for Interpretable and Controllable Neural Representations

  • 仅需每概念少于0.1%的标签,通过正则化引导关键概念定位。
  • 可逆调节或永久移除指定概念,对其他特征影响极小。
  • 适合需要可解释性与可控性的模型部署场景。

我们提出稀疏概念锚定方法,通过极少量监督(每概念标注少于0.1%样本)引导潜在空间中特定概念的位置,其余概念自由组织。训练结合激活归一化、分离正则化及锚点或子空间正则化,将稀有标注样本拉向预设方向或轴对齐子空间。该几何结构支持两种实用干预:推理时可逆地投影移除某概念分量,或通过针对性权重删减永久消除。在结构化自编码器上的实验表明,目标概念被选择性削弱,对正交特征影响微乎其微;完全移除后重建误差逼近理论下界。该方法为学习表示提供了可解释且可调控的行为路径。

原文摘要 · Abstract (English)

We introduce Sparse Concept Anchoring, a method that biases latent space to position a targeted subset of concepts while allowing others to self-organize, using only minimal supervision (labels for <0.1% of examples per anchored concept). Training combines activation normalization, a separation regularizer, and anchor or subspace regularizers that attract rare labeled examples to predefined directions or axis-aligned subspaces. The anchored geometry enables two practical interventions: reversible behavioral steering that projects out a concept's latent component at inference, and permanent removal via targeted weight ablation of anchored dimensions. Experiments on structured autoencoders show selective attenuation of targeted concepts with negligible impact on orthogonal features, and complete elimination with reconstruction error approaching theoretical bounds. Sparse Concept Anchoring therefore provides a practical pathway to interpretable, steerable behavior in learned representations.

可解释性概念控制稀疏监督潜在空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。