通过2万条真实对话,验证心理健康AI在真实场景下的安全性。
Beyond Simulations: What 20,000 Real Conversations Reveal About Mental Health AI Safety
- 用真实对话替代模拟测试,评估AI安全表现
- 自研系统在自杀等风险内容上错误率低于10%,远优于主流模型
- 临床审核发现仅3例未及时干预,证明系统整体可靠
心理健康AI安全通常依赖小规模仿真基准评估,可能无法反映实际部署中的语言与情境多样性。本研究将四个基准复现与真实世界对话生态审计结合,评估一款专用于心理健康领域的AI及六款前沿通用大模型(OpenAI GPT-5、GPT-5.1、GPT-5.2;DeepSeek V3;Google Gemini 3 Flash;Moonshot Kimi K2)。结果显示,该专用系统在自杀/自残、饮食障碍和物质滥用提示下的潜在有害内容率显著低于所有对比模型(CCDH基准:6.2% vs 18.0–52.0%,全部p < .001)。对2万条部署对话的审计中,临床审核确认无自杀风险对话遗漏危机资源,仅3例非自杀自伤(NSSI)对话未触发干预,内部条件率为3/800(0.38%)。为进一步推断群体水平表现,随机抽取600条对话进行盲评临床判定。结果显示,LLM判别器在设定阈值下敏感性达100%(6/6;95% CI 54.1–100%),特异性为99.2%(589/594),且所有6例经临床确认的高风险对话均成功提供危机资源,端到端零漏报(三倍上限95% CI 0.50%)。研究主张以生态审计补充预部署评估,提升心理健康AI安全性。
原文摘要 · Abstract (English)
Mental-health AI safety is typically evaluated with small, simulation-based benchmarks that may not reflect the linguistic and contextual diversity of deployment. We pair four benchmark replications with an ecological audit of real-world conversations to evaluate a purpose-built mental-health AI alongside six frontier general-purpose models spanning four families (OpenAI GPT-5, GPT-5.1, GPT-5.2; DeepSeek V3; Google Gemini 3 Flash; Moonshot Kimi K2). The purpose-built system produced significantly lower overall potentially harmful content rates than every frontier comparator on suicide/self-harm, eating-disorder, and substance-use prompts (CCDH Benchmark: Ash 6.2% vs frontier models 18.0-52.0%, all p < .001). In an audit of 20,000 deployment conversations, clinician review within the audit pipeline confirmed no suicide-risk conversations lacking crisis resources and three NSSI-related conversations without crisis intervention, a within-pipeline conditional rate of 3/800 (0.38%). To extend population-level inference, we adjudicated 600 conversations randomly sampled from the full deployment distribution via blinded clinician review. Against these labels, the LLM judge showed 100% sensitivity (6/6; 95% CI 54.1-100%) and 99.2% specificity (589/594) at its operating threshold, and the conversational model delivered crisis resources in all 6 clinician-confirmed cases (zero end-to-end false negatives; rule-of-three upper 95% CI 0.50%). Findings argue for ecological auditing as a complement to pre-deployment evaluation in mental-health AI safety
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。