arXiv:2606.21078cs.CL2026-06

用验证门控框架揭示大模型如何真正识别自杀意图

A Validation-Gated Mechanistic Account of Suicidality Detection in LLMs

  • 构建验证门控机制,仅在模型表现优于基线后才分析其内部特征
  • 发现中层特征具语义性且因果相关,跨模型与数据集稳定出现
  • 小模型已有表征但不行动,大模型才真正执行判断,适合临床研究参考

大语言模型在心理健康应用中被用于检测自杀内容,但其判断依据尚不明确。本文以自杀意图检测为案例,提出验证门控框架:仅当模型表现优于简单词法基线时,才允许对内部特征进行解释。该方法排除了在DeepSuiMind数据集上无法区分隐含自杀意图与普通痛苦的Llama-3.1-8B-Instruct模型。转向二分类任务后,发现中层特征具语义性、非关键词依赖、低秩且跨三个模型家族与三套自杀数据集重现;注册匹配控制(自杀对比抑郁)表明其更特异地追踪自杀风险。消融实验显示该特征对判断有因果影响,随机方向则无。可控性实验显示调整该特征可提升响应,但亦影响无关问题,故仅为必要非充分条件。最显著模式是编码与使用分离:小模型已具备表征能力,但仅大模型真正采取行动。证据基于英语Reddit文本,限制了临床适用性。

原文摘要 · Abstract (English)

Large language models are increasingly proposed for mental-health applications such as detecting suicidal content, raising the question of what they rely on. We study this mechanistically and use it to ask a narrower question: how to make a causal claim about a model's internal features more trustworthy. Our validation-gated framework, with suicidality detection as a case study, interprets a behavior only after the model is shown to perform it: a concept is admitted only once the model ranks it above a simple lexical baseline, and each subsequent property is tested against a matched control. This discipline yields negative as well as positive results. The gate rules out one task at the outset: on DeepSuiMind (Li et al. 2025), Llama-3.1-8B-Instruct cannot separate implicit suicidal intent from ordinary distress, so we do not analyze it. We turn to binary suicide detection, which it does perform. There we find a mid-network feature that appears semantic rather than keyword-based, is causally implicated in the decision (ablating it degrades the judgment; a random direction does not), is low-rank, and recurs across three model families and three suicide datasets. A register-matched control (suicide versus depression) suggests it tracks suicidality more specifically than general distress. Steering raises the model's response, but for unrelated questions too, so we treat it as necessary but not sufficient. The clearest pattern separates encoding from use: smaller models already represent suicidality, yet only larger ones appear to act on it. The positive evidence is English Reddit text, which limits the clinical reading.

自杀检测模型解释因果分析大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。