用三方反馈动态评估并优化AI心理咨询师,提升真实场景表现
Ψ-Arena: Interactive Assessment and Optimization of LLM-based Psychological Counselors with Tripartite Feedback
- 设计多阶段对话场景,用心理画像角色模拟真实咨询
- 从用户、咨询师、督导三视角评估,发现模型表现差异显著
- 闭环优化使性能最高提升141%,适合心理健康AI研发者
大语言模型在提供可扩展的心理健康支持方面展现出潜力,但评估其咨询能力对确保效果与安全至关重要。现有评估存在静态测试、单一视角和开环框架等局限。为此,我们提出Ψ-Arena,一个交互式框架,用于全面评估与优化基于LLM的咨询师,具备三大特点:(1) 真实场景交互,通过具有心理画像的非玩家角色(NPC)进行多阶段对话;(2) 三方评估,整合来访者、咨询师与督导的评价视角;(3) 闭环优化,利用诊断反馈迭代改进模型。在八种主流LLM上的实验表明,不同情境与视角下模型表现差异显著;基于反思的优化使咨询表现最高提升141%。我们期望PsychoArena成为推动可靠且符合人类价值的LLM在心理健康领域应用的基础资源。
原文摘要 · Abstract (English)
Large language models (LLMs) have shown promise in providing scalable mental health support, while evaluating their counseling capability remains crucial to ensure both efficacy and safety. Existing evaluations are limited by the static assessment that focuses on knowledge tests, the single perspective that centers on user experience, and the open-loop framework that lacks actionable feedback. To address these issues, we propose Ψ-Arena, an interactive framework for comprehensive assessment and optimization of LLM-based counselors, featuring three key characteristics: (1) Realistic arena interactions that simulate real-world counseling through multi-stage dialogues with psychologically profiled NPC clients, (2) Tripartite evaluation that integrates assessments from the client, counselor, and supervisor perspectives, and (3) Closed-loop optimization that iteratively improves LLM counselors using diagnostic feedback. Experiments across eight state-of-the-art LLMs show significant performance variations in different real-world scenarios and evaluation perspectives. Moreover, reflection-based optimization results in up to a 141% improvement in counseling performance. We hope PsychoArena provides a foundational resource for advancing reliable and human-aligned LLM applications in mental healthcare.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。