arXiv:2608.20345cs.CLcs.AI2026-08中稿 · ACM FAccT '26

AI心理助手难懂青少年的隐晦表达,误判风险高达34%。

When Vocabulary Comprehension Fails Clinical Reasoning: Evaluating Therapy Bots' Safety Risks for Generation Alpha

  • 构建青少年心理表达数据集与对话基准,测试大模型理解力
  • 模型词汇理解率76%-82%,但风险识别率仅64%-72%,差距达10-14个百分点
  • 发现六类误判模式,轻量缓解无效,需人工介入保障安全

对话式AI已成为生成代(2010-2024年出生)青少年非正式心理支持资源,美国13.1%的青少年(540万人)使用生成式AI获取心理健康建议。尽管这些系统基于大量心理学文献训练,但其对青少年特有的夸张语言、反讽积极、语义快速漂移和语境多义等特征的理解安全性尚未验证。在多起青少年因与AI聊天机器人互动导致死亡事件后,系统性评估至关重要。我们提出两个基准:(1)由母语者(ICC=0.72)和临床医生(kappa=0.78)验证的64条生成代心理表达;(2)75组多轮对话(共780轮),含标准/生成代双版本。评估了治疗类应用和通用聊天机器人的主流大模型架构——Claude、GPT-4o、Llama-3.1——结果显示,模型词汇理解率达76%-82%,但临床风险校准率仅为64%-72%,形成10-14百分点的词汇理解差距(p<.001,d>0.48),而人类治疗师仅3个百分点(p=.22)。该差距具有架构一致性,并随模糊性增加而扩大(7pp→18pp)。我们识别出六类失败模式:反讽掩盖(29pp)、轻描淡写接受(43pp)、非正式风格偏倚(24pp)、风险分层模糊(19pp)、语义漂移(19pp)、语境依赖暴力(7pp)。当三种及以上模式同时出现时,误判率高达94%。轻量级缓解措施无效;唯有重型结构化设计才能达到人类水平(成本提升6.4倍)。基于34%的基础误判率,估算每年约有146,880次危机被遗漏,我们建议强制采用人机协同架构、每季度开展青年专项验证、透明披露性能表现,并建立面向青少年的心理健康类AI监管框架。

原文摘要 · Abstract (English)

Conversational AI systems have become informal mental health support resources for Generation Alpha (Gen Alpha, born 2010-2024), with 13.1% of U.S. adolescents (5.4 million) using generative AI for mental health advice. While these systems, from therapy apps to general chatbots, rely on large language models trained on extensive psychological literature, their safety for youth communication patterns characterized by hyperbolic language, ironic positivity, rapid semantic drift, and contextual polysemy remains unvalidated. Following multiple adolescent deaths linked to AI chatbot interactions, systematic evaluation is critical. We present two benchmarks: (1) 64 Gen Alpha mental health expressions validated by native speakers (ICC=0.72) and clinicians (kappa=0.78); (2) 75 multi-turn conversations (780 turns) with paired Standard/Gen Alpha versions. Across evaluations of LLM architectures underlying therapy apps and general chatbots - Claude, GPT-4o, Llama-3.1 - models understand 76-82% of vocabulary but correctly calibrate only 64-72% of clinical risk, creating a 10-14 percentage point (pp) vocabulary-comprehension gap (p<.001, d>0.48) absent in human therapists (3pp, p=.22). The gap is architecturally consistent and widens with ambiguity (7pp -> 18pp). We identify six failure patterns: sarcasm masking (29pp), minimization acceptance (43pp), informal style bias (24pp), risk-stratified ambiguity (19pp), semantic drift (19pp), context-dependent violence (7pp). Patterns compound; three or more yield 94% miss rates. Lightweight mitigations fail; only heavy scaffolding achieves human performance (6.4x cost). With 34% baseline miss rate yielding 146,880 estimated annual missed crises, we recommend mandatory human-in-the-loop architectures, quarterly youth-specific validation, transparent performance disclosure, and regulatory frameworks for youth-facing mental health AI.

心理健康AI青少年风险识别大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。