arXiv:2502.17899cs.CL2025-02EMNLP被引 20

测试大模型识别隐性自杀倾向能力,发现普遍表现不佳。

Can Large Language Models Identify Implicit Suicidal Ideation? An Empirical Evaluation

  • 基于心理学框架构建1308条测试数据,评估模型识别隐性自杀意图能力。
  • 8个主流大模型在隐性自杀倾向检测中表现差,支持性回应也不够恰当。
  • 研究警示:当前大模型不适合直接用于心理危机干预,需更精细设计。

本文提出一个全面的评估框架,用于检验大型语言模型(LLMs)在自杀预防中的能力,重点关注两个关键方面:隐性自杀意念(IIS)的识别和适当支持性回应(PAS)的提供。我们引入 ewdata,一个基于心理框架(包括D/S-IAT和负性自动思维)并结合真实场景的1,308条测试案例的新数据集。通过对8个广泛使用的LLM在不同上下文设置下的大规模实验,发现当前模型在检测隐性自杀意念和支持性回应方面存在显著不足,凸显了将大模型应用于心理健康场景时的关键局限。研究强调,必须发展更复杂的方法来构建和评估适用于敏感心理应用的大模型。

原文摘要 · Abstract (English)

We present a comprehensive evaluation framework for assessing Large Language Models' (LLMs) capabilities in suicide prevention, focusing on two critical aspects: the Identification of Implicit Suicidal ideation (IIS) and the Provision of Appropriate Supportive responses (PAS). We introduce \ourdata, a novel dataset of 1,308 test cases built upon psychological frameworks including D/S-IAT and Negative Automatic Thinking, alongside real-world scenarios. Through extensive experiments with 8 widely used LLMs under different contextual settings, we find that current models struggle significantly with detecting implicit suicidal ideation and providing appropriate support, highlighting crucial limitations in applying LLMs to mental health contexts. Our findings underscore the need for more sophisticated approaches in developing and evaluating LLMs for sensitive psychological applications.

心理健康大模型评估自杀意念

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。