arXiv:2506.01329cs.CLcs.AI2025-06被引 6

用真实热线数据评估大模型危机识别能力,发现其表现接近专业人员。

Evaluating Large Language Models in Crisis Detection: A Real-World Benchmark from Psychological Support Hotlines

  • 构建540条真实心理热线语料库,评测大模型在四种危机任务上的表现。
  • 在自杀意图识别等任务上,大模型F1最高达0.907,部分任务优于人类。
  • 小模型经微调后反而超越大模型,提示高效训练策略的重要性。

心理援助热线是危机干预的重要渠道,但面临需求上升与资源有限的挑战。大语言模型(LLMs)在危机评估中具有潜力,但在情感敏感的真实临床场景中仍缺乏充分验证。本文提出PsyCrisisBench,基于杭州心理援助热线的540条标注话单,评估四大任务:情绪状态识别、自杀意念检测、自杀计划识别和风险评估。共测试64个来自15个模型家族的LLM(包括GPT、Claude、Gemini等闭源模型及Llama、Qwen、DeepSeek等开源模型),采用零样本、少样本和微调三种范式。结果显示,大模型在自杀意念检测(F1=0.880)、自杀计划识别(F1=0.779)和风险评估(F1=0.907)任务上表现优异,少样本提示和微调显著提升性能。相比受训人类操作员,大模型在自杀计划识别和风险评估上达到或超越人类水平,而人类在情绪状态识别和自杀意念检测上仍有优势。情绪状态识别难度较大(最高F1=0.709),可能因缺乏语音线索和语义模糊所致。值得注意的是,一个1.5B参数的微调模型(Qwen2.5-1.5B)在情绪和自杀意念任务上超越更大模型。结果表明,大模型在文本型危机评估中表现接近人类,具备互补优势。PsyCrisisBench为未来模型开发与临床伦理部署提供了可靠评估框架。

原文摘要 · Abstract (English)

Psychological support hotlines serve as critical lifelines for crisis intervention but encounter significant challenges due to rising demand and limited resources. Large language models (LLMs) offer potential support in crisis assessments, yet their effectiveness in emotionally sensitive, real-world clinical settings remains underexplored. We introduce PsyCrisisBench, a comprehensive benchmark of 540 annotated transcripts from the Hangzhou Psychological Assistance Hotline, assessing four key tasks: mood status recognition, suicidal ideation detection, suicide plan identification, and risk assessment. 64 LLMs across 15 model families (including closed-source such as GPT, Claude, Gemini and open-source such as Llama, Qwen, DeepSeek) were evaluated using zero-shot, few-shot, and fine-tuning paradigms. LLMs showed strong results in suicidal ideation detection (F1=0.880), suicide plan identification (F1=0.779), and risk assessment (F1=0.907), with notable gains from few-shot prompting and fine-tuning. Compared to trained human operators, LLMs achieved comparable or superior performance on suicide plan identification and risk assessment, while humans retained advantages on mood status recognition and suicidal ideation detection. Mood status recognition remained challenging (max F1=0.709), likely due to missing vocal cues and semantic ambiguity. Notably, a fine-tuned 1.5B-parameter model (Qwen2.5-1.5B) outperformed larger models on mood and suicidal ideation tasks. LLMs demonstrate performance broadly comparable to trained human operators in text-based crisis assessment, with complementary strengths across task types. PsyCrisisBench provides a robust, real-world evaluation framework to guide future model development and ethical deployment in clinical mental health.

大模型危机检测心理援助评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。