arXiv:2509.09712cs.CLcs.AI2025-09被引 2

用强化学习让小模型学会正念接纳疗法,效果比传统方法更好。

The Thinking Therapist: Training Large Language Models to Deliver Acceptance and Commitment Therapy using Supervised Fine-Tuning and Odds Ratio Policy Optimization

  • 采用概率优化训练,让模型掌握治疗过程而非简单模仿话术。
  • 使用专业评分器评估,新方法在治疗契合度上提升2.68分。
  • 思维链对模仿型训练有帮助,但对高级优化模型无额外增益。

接纳与承诺疗法(ACT)是新兴的第三波认知行为疗法,已在多种精神疾病中显示疗效。本研究探讨后训练方法与显式推理对小型开源大语言模型(LLM)执行ACT的影响。基于Mistral-Large生成的合成治疗对话,我们使用两种方法——监督微调(SFT)与几率比策略优化(ORPO),分别配合与不配合显式思维链(COT)推理步骤,对Llama-3.2-3b-Instruct进行训练。通过模拟治疗会话,由经人类评价微调过的LLM裁判,以ACT契合度量表(ACT-FM)和治疗师共情量表(TES)进行定量评估。结果显示,ORPO训练模型在ACT契合度(χ²(5) = 185.15, p < .001)和治疗共情(χ²(5) = 140.37, p < .001)上显著优于SFT及基础Instruct模型。思维链仅对SFT模型有效,平均提升ACT-FM得分2.68点(p < .001),对更优的ORPO或Instruct模型无显著作用。我们认为,ORPO的优势在于学习治疗‘过程’而非‘内容’,而思维链仅为模仿型训练提供必要支架。

原文摘要 · Abstract (English)

Acceptance and Commitment Therapy (ACT) is a third-wave cognitive behavioral therapy with emerging evidence of efficacy in several psychiatric conditions. This study investigates the impact of post-training methodology and explicit reasoning on the ability of a small open-weight large language model (LLM) to deliver ACT. Using synthetic ACT transcripts generated by Mistral-Large, we trained Llama-3.2-3b-Instruct with two distinct approaches, supervised fine-tuning (SFT) and odds ratio policy optimization (ORPO), each with and without an explicit chain-of-thought (COT) reasoning step. Performance was evaluated by comparing these four post-trained variants against the base Instruct model. These models were benchmarked in simulated therapy sessions, with performance quantitatively assessed on the ACT Fidelity Measure (ACT-FM) and the Therapist Empathy Scale (TES) by an LLM judge that had been fine-tuned on human evaluations. Our findings demonstrate that the ORPO-trained models significantly outperformed both their SFT and Instruct counterparts on ACT fidelity ($χ^2(5) = 185.15, p < .001$) and therapeutic empathy ($χ^2(5) = 140.37, p < .001$). The effect of COT was conditional as it provided a significant benefit to SFT models, improving ACT-FM scores by an average of 2.68 points ($p < .001$), while offering no discernible advantage to the superior ORPO or instruct-tuned variants. We posit that the superiority of ORPO stems from its ability to learn the therapeutic `process' over imitating `content,' a key aspect of ACT, while COT acts as a necessary scaffold for models trained only via imitation. This study establishes that preference-aligned policy optimization can effectively instill ACT competencies in small LLMs, and that the utility of explicit reasoning is highly dependent on the underlying training paradigm.

心理治疗强化学习大模型应用认知行为

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。