arXiv:2510.23845cs.CLcs.AI2025-10Conference of the …被引 7

构建首个含时间标签的多维度心理危机检测基准

CRADLE Bench: A Clinician-Annotated Benchmark for Multi-Faceted Mental Health Crisis and Safety Risk Detection

  • 基于临床标准设计七类危机场景,引入时间标注机制
  • 包含600条医生标注测试样本,420条开发样本,4000条自动标注训练数据
  • 提供不同共识标准下的模型,适合安全检测与医疗对话系统研究

检测自杀意念、性侵、家庭暴力、儿童虐待及性骚扰等心理健康危机,是语言模型面临的关键挑战。当用户与模型互动中出现此类情况时,模型需可靠识别,否则可能造成严重后果。本文提出CRADLE BENCH,一个面向多维度危机检测的基准。不同于以往仅覆盖少数危机类型的工作,本基准涵盖七类符合临床标准的危机类型,并首次引入时间标签。基准包含600条临床医生标注的评估样本和420条开发样本,另配有约4000条通过多模型多数投票集成自动标注的训练数据,其标注质量显著优于单模型标注。我们还基于共识与一致性的集成判断标准,对六种危机检测模型进行微调,提供不同标注标准下的互补模型。

原文摘要 · Abstract (English)

Detecting mental health crisis situations such as suicide ideation, rape, domestic violence, child abuse, and sexual harassment is a critical yet underexplored challenge for language models. When such situations arise during user--model interactions, models must reliably flag them, as failure to do so can have serious consequences. In this work, we introduce CRADLE BENCH, a benchmark for multi-faceted crisis detection. Unlike previous efforts that focus on a limited set of crisis types, our benchmark covers seven types defined in line with clinical standards and is the first to incorporate temporal labels. Our benchmark provides 600 clinician-annotated evaluation examples and 420 development examples, together with a training corpus of around 4K examples automatically labeled using a majority-vote ensemble of multiple language models, which significantly outperforms single-model annotation. We further fine-tune six crisis detection models on subsets defined by consensus and unanimous ensemble agreement, providing complementary models trained under different agreement criteria.

心理危机安全检测医疗对话

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。