arXiv:2412.00251cs.AIcs.HC2024-12被引 3

用小模型微调出能开展认知行为疗法的聊天助手,效果显著优于普通对话模型。

Fine-Tuning Open-Weight Language Models to Deliver Cognitive Behavioral Therapy for Depression: A Feasibility Study

  • 用合成心理咨询对话数据微调小规模大语言模型,专攻认知行为疗法。
  • 微调后模型在治疗量表上平均分提升11.33分,最优模型达67.86分。
  • 适合心理健康技术探索者,但需解决临床应用中的伦理与可靠性问题。

认知行为疗法(CBT)是治疗重度抑郁症的有效方法,但存在成本高、治疗师稀缺和污名化等障碍。本研究探索了微调小型开源大语言模型(LLMs)以提供CBT治疗的可行性。基于Nous Research对Llama 3.1 405b的微调生成的合成CBT对话数据,我们对Mistral 7b v0.3、Qwen 2.5 7b和Llama 3.1 8b三个模型进行微调。通过修改版认知治疗评分量表(CTRS)评估治疗契合度。所有微调模型与对应的指令微调版本对比,使用模拟患者对话进行测试,其中指令模型作为治疗师,DeepSeek-V2.5作为患者,由Gemini 1.5 Pro-002对模拟对话进行评分。结果显示,微调模型显著优于指令模型,总分平均提升11.33分(p < 0.001)。Llama 3.1 8b表现最佳(均值67.86 ± 7.24),其次为Qwen 2.5 7b(64.28 ± 9.55)和Mistral 7b v0.3(64.17 ± 9.79),差异具有统计学意义。微调模型能有效执行核心CBT技巧并表达共情,但在议程遵循、探索深度和长上下文连贯性方面仍存局限。研究证明,针对CBT的微调可在小型模型中有效编码治疗能力,但临床部署前仍需解决重大技术和伦理问题。

原文摘要 · Abstract (English)

Cognitive Behavioral Therapy (CBT) is a well-established, evidence-based treatment for Major Depressive Disorder. Unfortunately, there exist significant barriers to individuals accessing CBT, including cost, scarcity of therapists and stigma. This study explores the feasibility of fine-tuning small open weight large language models (LLMs) to deliver CBT for depression. Using synthetic CBT transcripts generated by the Nous Research fine-tune of Llama 3.1 405b, we fine-tuned three models: Mistral 7b v0.3, Qwen 2.5 7b, and Llama 3.1 8b. CBT fidelity was evaluated through a modified Cognitive Therapy Rating Scale (CTRS). All fine-tuned models were compared against each other, as well as their instruct-tuned variants. Simulated patient transcripts were generated for the purpose of evaluating model performance, with the instruct and CBT-tuned models acting as the therapist and DeepSeek-V2.5 acting as the patient. These simulated transcripts were evaluated on a modified CTRS by Gemini 1.5 Pro-002. Our findings demonstrated that the CBT-tuned models significantly outperformed their instruct-tuned counterparts, with an average improvement of 11.33 points (p < 0.001) on total CTRS score. Llama 3.1 8b had the strongest performance (mean CTRS score 67.86 +/- 7.24), followed by Qwen 2.5 7b (64.28 +/- 9.55) and Mistral 7b v0.3 (64.17 +/- 9.79), with these differences between models being statistically significant. The CBT-tuned models were competent in implementing core CBT techniques and providing empathetic responses, however, there were limitations observed in agenda adherence, exploration depth and long-context coherence. This study establishes that CBT specific fine-tuning can effectively encode therapeutic competencies in small LLMs, though significant technical and ethical considerations must be resolved prior to clinical deployment.

心理AI语言模型治疗模拟微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。