arXiv:2505.13480cs.CLcs.AI2025-05被引 2

用心理量表评估大模型识别自杀风险能力,表现接近人类。

Evaluating Reasoning LLMs for Suicide Screening with the Columbia-Suicide Severity Rating Scale

  • 用C-SSRS量表零样本测试6个大模型的自杀风险分级能力
  • Claude和GPT与人工标注高度一致,误判多在相邻等级间
  • 强调需人工监督,适合心理健康领域研究者参考

自杀预防仍是重大公共卫生挑战。尽管在线平台如Reddit的r/SuicideWatch曾为有自杀念头者提供表达与支持空间,但大语言模型(LLMs)的出现带来了新范式——个体可能开始向AI系统而非人类倾诉。本研究评估了六种模型(包括Claude、GPT、Mistral和LLaMA)使用哥伦比亚自杀严重度量表(C-SSRS)进行自动化自杀风险评估的能力。在7级严重度量表(0-6)上测试其零样本表现,结果表明Claude和GPT与人工标注高度一致,Mistral取得最低序数预测误差。多数模型表现出序数敏感性,误判主要发生在相邻等级之间。进一步分析了混淆模式、误判来源及伦理问题,强调需保持人工监督、提升透明度并谨慎部署。完整代码与补充材料见https://github.com/av9ash/llm_cssrs_code。

原文摘要 · Abstract (English)

Suicide prevention remains a critical public health challenge. While online platforms such as Reddit's r/SuicideWatch have historically provided spaces for individuals to express suicidal thoughts and seek community support, the advent of large language models (LLMs) introduces a new paradigm-where individuals may begin disclosing ideation to AI systems instead of humans. This study evaluates the capability of LLMs to perform automated suicide risk assessment using the Columbia-Suicide Severity Rating Scale (C-SSRS). We assess the zero-shot performance of six models-including Claude, GPT, Mistral, and LLaMA-in classifying posts across a 7-point severity scale (Levels 0-6). Results indicate that Claude and GPT closely align with human annotations, while Mistral achieves the lowest ordinal prediction error. Most models exhibit ordinal sensitivity, with misclassifications typically occurring between adjacent severity levels. We further analyze confusion patterns, misclassification sources, and ethical considerations, underscoring the importance of human oversight, transparency, and cautious deployment. Full code and supplementary materials are available at https://github.com/av9ash/llm_cssrs_code.

自杀风险大模型评估心理健康量表测评

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。