测试大模型是否具备心理咨询核心能力,发现其仍需改进。
Do Large Language Models Align with Core Mental Health Counseling Competencies?
- 构建基于NCMHCE的评测基准,评估22个大模型在五大能力上的表现。
- 顶尖模型在接诊与诊断上达标,但在共情和伦理实践上明显不足。
- 医疗类模型虽解释更优,但易出上下文错误,不适合独立使用。
大语言模型(LLMs)的快速发展为缓解全球心理健康专业人员短缺提供了可能,但其与核心心理咨询能力的契合度尚未充分研究。本文提出CounselingBench,一个基于NCMHCE的新型评测基准,用于评估22个通用及医学微调的大模型在五个关键能力维度上的表现。尽管前沿模型已达到最低胜任力阈值,但在专家级表现上仍有差距,尤其在核心咨询特质和职业实践与伦理方面表现较弱。值得注意的是,医学类大模型在准确性上并未优于通用模型,但提供更优的推理解释,同时产生更多与上下文相关错误。结果表明,开发适用于心理健康咨询的AI仍面临挑战,尤其在需要共情与精细推理的能力上。研究强调需针对核心咨询能力进行专门微调,并结合人工监督,才能安全投入实际应用。代码与数据见:https://github.com/cuongnguyenx/CounselingBench
原文摘要 · Abstract (English)
The rapid evolution of Large Language Models (LLMs) presents a promising solution to the global shortage of mental health professionals. However, their alignment with essential counseling competencies remains underexplored. We introduce CounselingBench, a novel NCMHCE-based benchmark evaluating 22 general-purpose and medical-finetuned LLMs across five key competencies. While frontier models surpass minimum aptitude thresholds, they fall short of expert-level performance, excelling in Intake, Assessment & Diagnosis but struggling with Core Counseling Attributes and Professional Practice & Ethics. Surprisingly, medical LLMs do not outperform generalist models in accuracy, though they provide slightly better justifications while making more context-related errors. These findings highlight the challenges of developing AI for mental health counseling, particularly in competencies requiring empathy and nuanced reasoning. Our results underscore the need for specialized, fine-tuned models aligned with core mental health counseling competencies and supported by human oversight before real-world deployment. Code and data associated with this manuscript can be found at: https://github.com/cuongnguyenx/CounselingBench
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。