arXiv:2601.01802cs.AI2026-01ACL被引 2

构建多阶段多疗法心理辅导AI评估基准,推动高真实感智能咨询发展。

PsychEval: A Multi-Session and Multi-Therapy Benchmark for High-Realism AI Psychological Counselor

  • 设计跨6-10阶段的多会话心理辅导任务,要求持续记忆与动态目标追踪。
  • 涵盖5种疗法及整合框架,覆盖6大核心主题,构建超2000个多样化来访者档案。
  • 提供18项疗法特异性与共享指标,支持临床级评估与自进化训练。

为开发可靠的AI心理评估系统,我们提出\texttt{PsychEval},一个跨多会话、多疗法且高度真实的基准测试。该基准解决三大挑战:1)能否训练出高真实感的AI咨询师?通过设计6-10会话的多阶段任务,要求具备记忆连续性、自适应推理与纵向规划能力,并标注超过677项元技能与4577项原子技能。2)如何训练多疗法AI咨询师?构建覆盖五类疗法(心理动力学、行为主义、认知行为疗法、人本存在主义、后现代主义)及整合疗法的数据集,采用统一三阶段临床框架,覆盖六大核心心理主题。3)如何系统评估AI咨询师?建立包含18项疗法特异与共用指标的综合评估体系,支撑客户端与咨询师端双重维度。同时构建超2000个多样化来访者画像。实验充分验证数据集的高质量与临床真实性。关键在于,\texttt{PsychEval}超越静态评估,可作为高保真强化学习环境,实现临床责任型自适应AI咨询师的自我演进训练。

原文摘要 · Abstract (English)

To develop a reliable AI for psychological assessment, we introduce \texttt{PsychEval}, a multi-session, multi-therapy, and highly realistic benchmark designed to address three key challenges: \textbf{1) Can we train a highly realistic AI counselor?} Realistic counseling is a longitudinal task requiring sustained memory and dynamic goal tracking. We propose a multi-session benchmark (spanning 6-10 sessions across three distinct stages) that demands critical capabilities such as memory continuity, adaptive reasoning, and longitudinal planning. The dataset is annotated with extensive professional skills, comprising over 677 meta-skills and 4577 atomic skills. \textbf{2) How to train a multi-therapy AI counselor?} While existing models often focus on a single therapy, complex cases frequently require flexible strategies among various therapies. We construct a diverse dataset covering five therapeutic modalities (Psychodynamic, Behaviorism, CBT, Humanistic Existentialist, and Postmodernist) alongside an integrative therapy with a unified three-stage clinical framework across six core psychological topics. \textbf{3) How to systematically evaluate an AI counselor?} We establish a holistic evaluation framework with 18 therapy-specific and therapy-shared metrics across Client-Level and Counselor-Level dimensions. To support this, we also construct over 2,000 diverse client profiles. Extensive experimental analysis fully validates the superior quality and clinical fidelity of our dataset. Crucially, \texttt{PsychEval} transcends static benchmarking to serve as a high-fidelity reinforcement learning environment that enables the self-evolutionary training of clinically responsible and adaptive AI counselors.

心理AI多疗法评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。