用问卷生成心理治疗对话,保护隐私还能提升模型表现
Roleplaying with Structure: Synthetic Therapist-Client Conversation Generation from Questionnaires
- 基于真实问卷构建结构化对话生成流程,不泄露敏感信息
- 专家评估显示生成对话更贴近真实治疗场景,优于现有模型
- 释放数据集与微调模型,适合心理AI研究者使用
大型语言模型(LLMs)在心理健康领域合成数据生成方面具有潜力。但以往工作受限于隐私政策,主要依赖通用信息。本文提出全面的合成心理治疗对话语料库,构建了基于结构化问卷的心理治疗生成管道SQPsych(Structured Questionnaire-based Psychotherapy),利用真实客户档案与心理问卷生成数据,确保不泄露敏感信息。我们对多种开源权重的LLM在生成语料SQPsychConv上进行微调,并通过自动评估与训练过的心理治疗师人工评价进行测试。结果表明,标准基准无法充分反映该数据集的优势,但专家判断显示,使用我们的方法使模型在扮演治疗师方面显著提升;专家一致更偏好由本模型生成的治疗会话,优于其他面向心理健康领域的LLM。代码、微调模型SQPsychLLM及语料库已公开于https://ai-mh.github.io/SQPsych.html。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are promising tools for synthetic data generation in mental health. However, privacy policies and restrictions forced previous work to rely mainly on generic information. We present a comprehensive corpus of synthetic therapist-client conversations generated through LLMs. We construct our generation pipeline, SQPsych (Structured Questionnaire-based Psychotherapy), which uses real structured client profiles and psychological questionnaires without leaking any sensitive data. We fine-tune various open-weight LLMs on our generated corpus, SQPsychConv , and test them through both automatic benchmarks and human evaluation with trained psychotherapists. We find that standard benchmarks do not adequately capture the strengths of our dataset, but expert judgment shows that SQPsych makes LLMs significantly better at therapist roleplaying. Experts also consistently prefer therapy sessions generated by our models compared to other mental-health-oriented LLMs. We release our code, fine-tuned models SQPsychLLM, and corpora at https://ai-mh.github.io/SQPsych.html.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。