arXiv:2503.03601cs.CLcs.IT2025-03ACL被引 14

用稀疏自编码器解析大模型文本的写作特征,揭示其与人类文本的本质差异。

Feature-Level Insights into Artificial Text Detection with Sparse Autoencoders

  • 通过稀疏自编码器提取模型残差流中的可解释特征
  • 发现现代大模型在信息密集领域有独特写作风格
  • 适合关注AI文本检测与可解释性的研究者

随着大语言模型(LLM)的兴起,人工文本检测(ATD)变得日益重要。尽管已有诸多努力,但尚无单一算法能在不同类型的未见文本上保持稳定表现,也无法保证对新大模型的有效泛化。可解释性在此目标中起关键作用。本研究利用稀疏自编码器(SAE)从Gemma-2-2b的残差流中提取特征,识别出既可解释又高效的特征。通过领域和模型特定统计、引导方法以及人工或基于LLM的分析,我们深入理解了不同模型生成文本与人类文本的差异。结果表明,即使在个性化提示下能生成类人输出,现代大模型在信息密集领域仍表现出独特的写作风格。

原文摘要 · Abstract (English)

Artificial Text Detection (ATD) is becoming increasingly important with the rise of advanced Large Language Models (LLMs). Despite numerous efforts, no single algorithm performs consistently well across different types of unseen text or guarantees effective generalization to new LLMs. Interpretability plays a crucial role in achieving this goal. In this study, we enhance ATD interpretability by using Sparse Autoencoders (SAE) to extract features from Gemma-2-2b residual stream. We identify both interpretable and efficient features, analyzing their semantics and relevance through domain- and model-specific statistics, a steering approach, and manual or LLM-based interpretation. Our methods offer valuable insights into how texts from various models differ from human-written content. We show that modern LLMs have a distinct writing style, especially in information-dense domains, even though they can produce human-like outputs with personalized prompts.

文本检测可解释性大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。