arXiv:2505.15038cs.CLcs.AI2025-05Conference of the …被引 13

用稀疏自编码器清除语言模型控制向量中的噪声,提升控制精度。

Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering

  • 通过选择性保留最区分度高的编码潜变量,重构隐藏表示。
  • 在6个挑战性概念上,控制成功率提升4-16%。
  • 适合需要精准可控的生成任务研究者使用。

线性概念向量能有效操控大语言模型,但现有方法在多样化数据集中受噪声特征影响,降低控制鲁棒性。本文提出稀疏自编码器去噪概念向量(SDCV),通过选择性保留最具区分性的SAE潜变量,并重构隐藏表示来分离信号与噪声。关键思路是放大表现最佳的top-k潜变量激活值,以明确区分正负样本。应用于线性探测和均值差异法时,SDCV在六个挑战性概念上均稳定提升控制成功率4-16%,同时保持主题相关性。

原文摘要 · Abstract (English)

Linear concept vectors effectively steer LLMs, but existing methods suffer from noisy features in diverse datasets that undermine steering robustness. We propose Sparse Autoencoder-Denoised Concept Vectors (SDCV), which selectively keep the most discriminative SAE latents while reconstructing hidden representations. Our key insight is that concept-relevant signals can be explicitly separated from dataset noise by scaling up activations of top-k latents that best differentiate positive and negative samples. Applied to linear probing and difference-in-mean, SDCV consistently improves steering success rates by 4-16\% across six challenging concepts, while maintaining topic relevance.

语言模型概念控制去噪自编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。