用100位专家评估大模型在心理问答中的表现,发现其常给出不安全建议。
CounselBench: A Large-Scale Expert Evaluation and Adversarial Benchmarking of Large Language Models in Mental Health Question Answering
- 邀请100位心理健康专家评估4个大模型对真实患者提问的回答
- 模型在6个维度得分尚可,但普遍存在过度泛化和危险建议问题
- 专设对抗性测试集,暴露不同模型的典型错误模式
医疗问答基准多聚焦选择题或事实型任务,对患者真实开放问题的回答研究不足。这一缺口在心理健康领域尤为严重,因患者提问常混合症状、治疗担忧与情感需求,需兼顾临床谨慎与情境敏感。我们提出CounselBench,一个由100位心理健康专业人士参与构建的大规模基准,用于评估和压力测试大语言模型(LLMs)在真实求助场景下的表现。第一部分CounselBench-EVAL包含对GPT-4、LLaMA 3、Gemini及在线人类治疗师在公共论坛CounselChat上回答的2,000条答案的专家评估,每条回答在六个临床相关维度打分,并附有片段级标注与书面理由。专家评估显示,尽管模型在多个维度表现良好,但仍反复出现非建设性反馈、过度泛化、个性化与相关性不足等问题,且频繁被标记存在安全风险,尤其是未经授权的医疗建议。后续实验表明,模型评分系统会系统性高估响应质量,忽视人类专家指出的安全隐患。为更直接探查失败模式,我们构建了包含120个专家撰写的心理健康问题的对抗性数据集CounselBench-Adv。对九个大模型生成的1,080条回应进行专家评估,揭示出一致且模型特有的故障模式。CounselBench共同建立了一个基于临床实践的框架,用于评估大模型在心理健康问答中的表现。
原文摘要 · Abstract (English)
Medical question answering (QA) benchmarks often focus on multiple-choice or fact-based tasks, leaving open-ended answers to real patient questions underexplored. This gap is particularly critical in mental health, where patient questions often mix symptoms, treatment concerns, and emotional needs, requiring answers that balance clinical caution with contextual sensitivity. We present CounselBench, a large-scale benchmark developed with 100 mental health professionals to evaluate and stress-test large language models (LLMs) in realistic help-seeking scenarios. The first component, CounselBench-EVAL, contains 2,000 expert evaluations of answers from GPT-4, LLaMA 3, Gemini, and online human therapists on patient questions from the public forum CounselChat. Each answer is rated across six clinically grounded dimensions, with span-level annotations and written rationales. Expert evaluations show that while LLMs achieve high scores on several dimensions, they also exhibit recurring issues, including unconstructive feedback, overgeneralization, and limited personalization or relevance. Responses were frequently flagged for safety risks, most notably unauthorized medical advice. Follow-up experiments show that LLM judges systematically overrate model responses and overlook safety concerns identified by human experts. To probe failure modes more directly, we construct CounselBench-Adv, an adversarial dataset of 120 expert-authored mental health questions designed to trigger specific model issues. Expert evaluation of 1,080 responses from nine LLMs reveals consistent, model-specific failure patterns. Together, CounselBench establishes a clinically grounded framework for benchmarking LLMs in mental health QA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。