arXiv:2606.10380cs.CLcs.AI2026-06

构建对话式心理危机检测数据集,提升早期风险识别能力

Expert-Level Crisis Detection in Mental Health Conversations

论文配图:Expert-Level Crisis Detection in Mental Health Conversations
图 1 · 摘自论文原文
  • 构建多轮对话中的逐轮危机检测基准,标注600段真实对话
  • 模型识别风险出现时机的准确率仅40%-60%,远低于识别风险存在
  • 提出预警-确认评估协议,更贴近临床实际干预需求

现实中的危机干预本质上是对话过程,但现有研究多聚焦静态文本。当应用于多轮对话时,当前模型性能显著下降,难以追踪随上下文演变而浮现的风险信号。为此,我们提出CRADLE-Dialogue,一个由临床医生标注的多轮对话中逐轮危机检测基准。该数据集包含600段对话,对自杀意念、自残、儿童虐待等临床风险进行多标签标注,并区分过去与当前风险状态。我们进一步提出一种预警-确认评估协议,将早期预警信号(Alert)与危机明确显现的回合(Confirm)区分开来,反映临床干预需在风险显性化前介入的需求。实验表明,识别风险何时出现比确认其存在更难,模型在微平均F1上仅为中等40%至高60%。此外,我们发布了合成训练语料和一个320亿参数模型,显著优于现有开源模型,在逐轮、对话级及仅确认项评估中达到或超越专有模型水平。

原文摘要 · Abstract (English)

Real-world crisis intervention is inherently conversational, yet existing research largely focuses on static texts.Real-world crisis intervention is inherently conversational, yet existing research largely focuses on static texts. When applied to multi-turn dialogues, current models exhibit significant performance degradation, struggling to track risk signals that emerge as context evolves. To address this gap, we introduce CRADLE-Dialogue, a clinician-annotated benchmark for turn-level crisis detection in conversational settings. The dataset features 600 dialogues with multi-label annotations across clinically grounded risks, including suicide ideation, self-harm, and child abuse, distinguishing past from ongoing risk. We further propose an Alert-Confirm evaluation protocol that distinguishes early warning signals (Alert) from turns where a specific crisis becomes explicitly identifiable (Confirm), reflecting the clinical need to intervene before risk becomes explicit. Experiments show that identifying when risk emerges is much harder than recognizing that it exists: models achieve only mid-40% to high-60% Micro F1. Additionally, we release a synthetic training corpus and a 32B-parameter model that substantially outperforms existing open-source models and achieves competitive or superior results against proprietary models across turn-level, dialogue-level, and confirm-only evaluation settings.

心理危机检测对话理解多轮对话临床应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。