用合成数据训练隐私保护的心理健康大模型,避免真实对话泄露风险。
Towards Privacy-Preserving Mental Health Support with Large Language Models
- 通过多智能体角色扮演生成高质量咨询对话数据
- 合成数据使模型在测评中表现媲美现有专业模型
- 采用联邦学习与差分隐私,降低数据泄露风险,适合医疗AI研究者
大型语言模型在心理健康支持方面展现出潜力,但其训练受限于真实咨询对话的稀缺性与敏感性。本文提出面向隐私保护的心理健康大模型 MindChat 及其配套数据集 MindCorpus,该数据集通过多智能体角色扮演框架构建。为生成高质量对话,系统采用双闭环反馈机制:一是逐轮批判与修订以提升会话连贯性与咨询适宜性;二是会话级策略优化以逐步丰富咨询师行为。为应对分布式数据所有权带来的隐私风险,采用参数高效且可微分的 LoRA 适配器进行联邦微调,并结合差分隐私优化,降低成员推断与记忆攻击风险。在合成数据质量评估与咨询能力测试中,实验表明 MindCorpus 显著提升训练效果,而 MindChat 在自动评测与人工评估下均优于或媲美现有通用及专业模型,同时在成员推断攻击下表现出更少的隐私泄漏。
原文摘要 · Abstract (English)
Large language models (LLMs) have shown promise for mental health support, yet training such models is constrained by the scarcity and sensitivity of real counseling dialogues. In this article, we present MindChat, a privacy-preserving LLM for mental health support, together with MindCorpus, a synthetic multi-turn counseling dataset constructed via a multi-agent role-playing framework. To synthesize high-quality counseling data, the developed dialogue-construction framework employs a dual closed-loop feedback design to integrate psychological expertise and counseling techniques through role-playing: (i) turn-level critique-and-revision to improve coherence and counseling appropriateness within a session, and (ii) session-level strategy refinement to progressively enrich counselor behaviors across sessions. To mitigate privacy risks under decentralized data ownership, we fine-tune the base model using federated learning with parameter-efficient LoRA adapters and incorporate differentially private optimization to reduce membership and memorization risks. Experiments on synthetic-data quality assessment and counseling capability evaluation show that MindCorpus improves training effectiveness and that MindChat is competitive with existing general and counseling-oriented LLM baselines under both automatic LLM-judge and human evaluation protocols, while exhibiting reduced privacy leakage under membership inference attacks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。