首个融合共情、专业与推理的中文心理大模型,提升心理问答准确性。
Psyche-R1: Towards Reliable Psychological LLMs through Unified Empathy, Expertise, and Reasoning
- 构建统一框架,整合共情对话、专业知识与逻辑推理能力。
- 生成超7.5万条带详细推理的心理问题数据,7.3万条共情对话。
- 7B小模型性能媲美671B大模型,适合心理辅助应用落地。
在心理健康专业人才短缺的背景下,将大语言模型(LLMs)应用于心理服务具有巨大潜力。现有研究多聚焦于情感支持与共情对话,缺乏对推理机制的重视。为此,本文提出首个中文心理大模型 extit{Psyche-R1},首次实现共情、心理专业知识与推理能力的统一集成。基于创新的数据构建流程,我们生成了超过75,000条高质量心理问题及其详细推理过程,通过迭代提示-推理优化实现;同时收集了73,000条共情对话。采用混合训练策略:利用多模型交叉选择识别困难样本,通过组相对策略优化(GRPO)增强推理能力;其余数据用于监督微调(SFT),提升共情回复与领域知识。大量实验表明, extit{Psyche-R1} 在多个心理基准测试中表现优异,其7B版本性能媲美671B的 exttt{DeepSeek-R1}。
原文摘要 · Abstract (English)
Amidst a shortage of qualified mental health professionals, the integration of large language models (LLMs) into psychological applications offers a promising way to alleviate the growing burden of mental health disorders. Recent reasoning-augmented LLMs have achieved remarkable performance in mathematics and programming, while research in the psychological domain has predominantly emphasized emotional support and empathetic dialogue, with limited attention to reasoning mechanisms that are beneficial to generating accurate responses. Therefore, in this paper, we propose \logopsyche\textit{Psyche-R1}, the first Chinese psychological LLM that jointly integrates empathy, psychological expertise, and reasoning, built upon a novel data curation pipeline. Specifically, we design a comprehensive data synthesis pipeline that produces over 75k high-quality psychological questions paired with detailed rationales, generated through an iterative prompt-rationale optimization procedure, along with 73k empathetic dialogues. Subsequently, we employ a hybrid training strategy wherein challenging samples are identified through a multi-LLM cross-selection strategy for group relative policy optimization (GRPO) to improve reasoning ability, while the remaining data are used for supervised fine-tuning (SFT) to enhance empathetic response generation and psychological domain knowledge. Extensive experiment results demonstrate the effectiveness of \textit{Psyche-R1} across several psychological benchmarks, where our 7B \textit{Psyche-R1} achieves comparable results to 671B \texttt{DeepSeek-R1}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。