发现并验证了大模型对齐训练带来的AI风格特征,可定位并消除。
Measuring, Localizing, and Ablating Alignment Signatures in LLMs

- 通过对比生成文本,发现对齐训练使输出更像AI
- 提出PASTA方法,能有效降低检测率15%-30%以上
- 适合关注模型可解释性与规避AI检测的研究者
对齐语言模型常表现出可识别的AI风格,但其与后训练过程及内部表征的关系仍不明确。本文比较人类文本、基础模型生成与对齐模型生成在相同人类源前缀下的表现。结果表明,对齐生成文本的人类语料亲和度更低,AI检测率更高,说明后训练使生成内容偏离人类风格,趋向检测可见的AI化文本。为此,我们提出PASTA(Post-training Alignment Signature Targeted Ablation)——一种无需再训练的方法,通过分析对齐模型与基础模型的残差差异,估计后训练对齐特征方向,并在解码中剔除该方向。在11个对齐模型和6种AI检测器上,PASTA显著降低多数模型的检测率(平均下降15%-30%),且跨检测器效果良好,随机方向无法复现此效果。定性分析显示,去除该方向后生成文本仍保持相关性与连贯性,但风格多样性提升。结果表明,后训练引发的AI风格效应可被测量、定位并因果验证。
原文摘要 · Abstract (English)
Aligned language models often exhibit a recognizable AI-like style, yet its connection to post-training and internal representations remains poorly understood. In this work, we study whether post-training introduces or amplifies AI-like stylistic regularities and whether these regularities have a localized internal signature. To this end, we compare human text, base-model generations, and aligned-model generations under matched human-source prefixes. Aligned generations show lower human-corpus affinity and higher AI-detection rates than base generations, suggesting that post-training shifts generated text away from human-corpus style and toward detector-visible AI-like text. We then introduce PASTA (Post-training Alignment Signature Targeted Ablation), a training-free method that estimates a post-training alignment signature from aligned-base residual contrasts and ablates the corresponding direction during decoding. Across 11 aligned models and 6 AI detectors, PASTA lowers the detection rate for most aligned models; this effect transfers well across detectors and is not reproduced by random directions. Qualitative analysis suggests that PASTA generations remain relevant and coherent while exhibiting greater stylistic variation. Together, these results show that AI-like stylistic effects of post-training can be measured, localized, and causally tested through activation ablation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。