用专业心理咨询标准训练模型,让大模型更懂心理疏导。
Preference Learning Unlocks LLMs' Psycho-Counseling Skills
- 构建36000对高质量咨询回复偏好数据集
- 模型在对话中胜过GPT-4o达87%胜率
- 适合心理AI研究者和临床辅助系统开发者
将大语言模型应用于心理辅导是缓解心理健康服务供需差距的新兴方向。但现有模型难以持续生成有效回应,主要因缺乏高质量真实咨询数据,且其内容受隐私保护难以获取。同时,可用数据中治疗师回复质量参差不齐,难以评估。本文提出一套专业全面的评估准则,基于该准则构建包含3.6万组偏好对比对的PsychoCounsel-Preference数据集,其标注符合专业治疗师偏好,为评估与优化模型提供坚实基础。实验表明,该数据集在奖励建模与偏好学习中表现优异,最佳模型PsychoCounsel-Llama3-8B在与GPT-4o的对比中取得87%胜率。相关数据集、模型及奖励模型已开源,地址:https://hf.co/Psychotherapy-LLM。
原文摘要 · Abstract (English)
Applying large language models (LLMs) to assist in psycho-counseling is an emerging and meaningful approach, driven by the significant gap between patient needs and the availability of mental health support. However, current LLMs struggle to consistently provide effective responses to client speeches, largely due to the lack of supervision from high-quality real psycho-counseling data, whose content is typically inaccessible due to client privacy concerns. Furthermore, the quality of therapists' responses in available sessions can vary significantly based on their professional training and experience. Assessing the quality of therapists' responses remains an open challenge. In this work, we address these challenges by first proposing a set of professional and comprehensive principles to evaluate therapists' responses to client speeches. Using these principles, we create a preference dataset, PsychoCounsel-Preference, which contains 36k high-quality preference comparison pairs. This dataset aligns with the preferences of professional psychotherapists, providing a robust foundation for evaluating and improving LLMs in psycho-counseling. Experiments on reward modeling and preference learning demonstrate that PsychoCounsel-Preference is an excellent resource for LLMs to acquire essential skills for responding to clients in a counseling session. Our best-aligned model, PsychoCounsel-Llama3-8B, achieves an impressive win rate of 87% against GPT-4o. We release PsychoCounsel-Preference, PsychoCounsel-Llama3-8B and the reward model PsychoCounsel Llama3-8B-Reward to facilitate the research of psycho-counseling with LLMs at: https://hf.co/Psychotherapy-LLM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。