用小模型模拟真实抑郁患者,效果比GPT-4还强
Eeyore: Realistic Depression Simulation via Supervised and Preference Optimization
- 通过真实对话数据构建抑郁特征库,指导模型微调
- 分两阶段优化:自动生成偏好+专家校准,提升真实感
- 适合心理治疗培训、临床角色扮演等需要真实患者模拟的场景
大型语言模型虽被用于心理健康训练与患者模拟,但仍难以真实呈现多样化的个体特征和心理状态。我们提出Eeyore,一个80亿参数的模型,通过结构化对齐框架实现更真实的抑郁模拟,全程融入领域专家意见。首先,系统收集真实抑郁相关对话,提取抑郁特征以指导数据筛选和心理画像构建,并用该数据集对Eeyore进行指令微调,确保角色符合心理档案。其次,为增强真实感,采用迭代式偏好优化:先使用模型自生成偏好,再结合少量专家标注偏好进行校准。整个流程中,我们与领域专家协作,开发交互界面验证特征提取并持续优化心理画像,实现临床上有意义的角色定制。尽管模型规模较小,Eeyore在语言真实性与心理档案一致性上均超越使用SOTA提示策略的GPT-4o。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have been previously explored for mental healthcare training and therapy client simulation, but they still fall short in authentically capturing diverse client traits and psychological conditions. We introduce \textbf{Eeyore}, an 8B model optimized for realistic depression simulation through a structured alignment framework, incorporating expert input at every stage. First, we systematically curate real-world depression-related conversations, extracting depressive traits to guide data filtering and psychological profile construction, and use this dataset to instruction-tune Eeyore for profile adherence. Next, to further enhance realism, Eeyore undergoes iterative preference optimization -- first leveraging model-generated preferences and then calibrating with a small set of expert-annotated preferences. Throughout the entire pipeline, we actively collaborate with domain experts, developing interactive interfaces to validate trait extraction and iteratively refine structured psychological profiles for clinically meaningful role-play customization. Despite its smaller model size, the Eeyore depression simulation outperforms GPT-4o with SOTA prompting strategies, both in linguistic authenticity and profile adherence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。