arXiv:2604.02359cs.CLcs.AI2026-04被引 4

用大模型当裁判评估精神分裂患者对话安全,效果接近医生共识。

Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis

论文配图:Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis
图 1 · 摘自论文原文
  • 设计7项临床专家制定的安全标准,针对精神分裂患者对话风险。
  • 大模型裁判与人类共识一致度达0.56-0.75,最佳单模型胜过多人投票。
  • 为心理健康领域大模型安全评估提供可扩展的临床验证方法。

通用大语言模型在心理健康支持中日益普及,但高频使用可能加剧精神分裂症患者的妄想和幻觉。现有评估缺乏临床验证且难以扩展。本研究聚焦精神分裂症,构建了7项临床专家定义的安全准则,建立人工共识数据集,并测试以大模型为评估者(LLM-as-a-Judge)或多个模型投票(LLM-as-a-Jury)的自动化评估方式。结果表明,大模型裁判与人类共识一致性良好(Cohen's $κ_{\text{human} \times \text{gemini}} = 0.75$, $κ_{\text{human} \times \text{qwen}} = 0.68$, $κ_{\text{human} \times \text{kimi}} = 0.56$),最优单模型表现略优于多模型投票($κ_{\text{human} \times \text{jury}} = 0.74$)。研究为心理健康场景下大模型安全评估提供了兼具临床基础与可扩展性的新路径。

原文摘要 · Abstract (English)

General-purpose Large Language Models (LLMs) are becoming widely adopted by people for mental health support. Yet emerging evidence suggests there are significant risks associated with high-frequency use, particularly for individuals suffering from psychosis, as LLMs may reinforce delusions and hallucinations. Existing evaluations of LLMs in mental health contexts are limited by a lack of clinical validation and scalability of assessment. To address these issues, this research focuses on psychosis as a critical condition for LLM safety evaluation by (1) developing and validating seven clinician-informed safety criteria, (2) constructing a human-consensus dataset, and (3) testing automated assessment using an LLM as an evaluator (LLM-as-a-Judge) or taking the majority vote of several LLM judges (LLM-as-a-Jury). Results indicate that LLM-as-a-Judge aligns closely with the human consensus (Cohen's $κ_{\text{human} \times \text{gemini}} = 0.75$, $κ_{\text{human} \times \text{qwen}} = 0.68$, $κ_{\text{human} \times \text{kimi}} = 0.56$) and that the best judge slightly outperforms LLM-as-a-Jury (Cohen's $κ_{\text{human} \times \text{jury}} = 0.74$). Overall, these findings have promising implications for clinically grounded, scalable methods in LLM safety evaluations for mental health contexts.

大模型安全精神健康自动评估临床验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。