arXiv:2607.21988cs.CL2026-07

分析大模型如何表征自残内容,揭示其深层表示规律。

Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study

论文配图:Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study
图 1 · 摘自论文原文
  • 在多个模型中探测自残信息的表征位置,发现集中在最后3%-7%层。
  • 自残语义方向的线性可分性与检测准确率不完全相关。
  • Gemma-3-4B表现独特,以更复杂方式编码自残语义。

自残内容在自然语言处理中极难检测,且属于高风险任务,需极高准确率以实现及时干预或标记潜在风险用户。本文分析大模型对自残内容的表征方式,为自残检测、模型干预与治理提供支持。研究聚焦两个数据集(X-Sensitive 和 SH-Detection)及四类模型,开展两项实验:(1) 在各模型所有层上训练并评估线性探测器,结果显示自残信息在两数据集上均集中于网络最后93%至97%深度(即最后3%-7%层);(2) 提取对比性的自残方向,经归一化后发现,最准确的探测器并非最具线性可分性。尤其发现,Gemma-3-4B以略有不同的、更复杂的机制表示该对比方向。

原文摘要 · Abstract (English)

Self-harm content is particularly challenging to detect using NLP techniques, and is also a high-stakes task which requires the highest accuracy to enable timely intervention or flagging at-risk users. We therefore present an analysis of how LLMs represent such self-harm content, which has downstream applications in self-harm detection, LLM intervention and governance and policing. In this paper, we focus on two datasets and four models, and perform two main experiments: (1) We train and evaluate linear probes across all layers of each model on two self-harm datasets: X-Sensitive and SH-Detection. Across both corpora, self-harm information crystallizes in the final 3 - 7% of network layers (93 to 97% depth). (2) We extract contrastive self-harm directions and, after performing a normaliation step, we find that the most accurate probes are not necessarily the most linearly separable. In particular, we find Gemma-3-4B to represent this \textit{contrastive self-harm direction} in a slightly different, more intricate way than the other LLMs.

自残检测大模型分析表示学习语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。