arXiv:2510.01611cs.AIcs.CL2025-10被引 2

评测大模型能否通过心理咨询师执照考试,发现只有顶尖模型达标。

PsychCounsel-Bench: Evaluating the Psychology Intelligence of Large Language Models

  • 构建基于美国心理咨询师执照考题的测评基准PsychCounsel-Bench
  • GPT-4o等前沿模型通过率超70%门槛,小模型远未达标
  • 为心理类大模型开发提供可量化的评估标准,适合研究者与开发者

大语言模型在生成任务中表现卓越,但在需要认知能力的应用如心理辅导方面潜力尚未释放。本文核心问题为:大模型能否有效用于心理辅导?为验证其胜任力,需考察其是否能通过美国国家心理咨询师认证考试(NCE),该考试要求约70%正确率方可通过。为此,我们提出PsychCounsel-Bench,一个基于美国国家心理咨询师考试的基准,包含约2,252道精心设计的单选题,覆盖心理学多个子领域,要求深度理解。评估显示,GPT-4o、Llama3.3-70B和Gemma3-27B等先进模型表现优异,远超及格线;而Qwen2.5-7B、Mistral-7B等开源小模型则显著低于及格线。结果表明,仅前沿大模型具备达到心理咨询标准的能力,凸显心理导向大模型发展的机遇与挑战。数据集已公开:https://github.com/cloversjtu/PsychCounsel-Bench

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated remarkable success across a wide range of industries, primarily due to their impressive generative abilities. Yet, their potential in applications requiring cognitive abilities, such as psychological counseling, remains largely untapped. This paper investigates the key question: \textit{Can LLMs be effectively applied to psychological counseling?} To determine whether an LLM can effectively take on the role of a psychological counselor, the first step is to assess whether it meets the qualifications required for such a role, namely the ability to pass the U.S. National Counselor Certification Exam (NCE). This is because, just as a human counselor must pass a certification exam to practice, an LLM must demonstrate sufficient psychological knowledge to meet the standards required for such a role. To address this, we introduce PsychCounsel-Bench, a benchmark grounded in U.S.national counselor examinations, a licensure test for professional counselors that requires about 70\% accuracy to pass. PsychCounsel-Bench comprises approximately 2,252 carefully curated single-choice questions, crafted to require deep understanding and broad enough to cover various sub-disciplines of psychology. This benchmark provides a comprehensive assessment of an LLM's ability to function as a counselor. Our evaluation shows that advanced models such as GPT-4o, Llama3.3-70B, and Gemma3-27B achieve well above the passing threshold, while smaller open-source models (e.g., Qwen2.5-7B, Mistral-7B) remain far below it. These results suggest that only frontier LLMs are currently capable of meeting counseling exam standards, highlighting both the promise and the challenges of developing psychology-oriented LLMs. We release the proposed dataset for public use: https://github.com/cloversjtu/PsychCounsel-Bench

心理模型测评基准大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。