用强化学习让对话有治疗阶段目标,提升心理咨询的结构化水平。
DeepSAGE: Stage-Aware Reinforcement Learning for Structured CBT Counseling Dialogue
- 将认知行为疗法首阶段拆为11个治疗阶段,由外部控制器判断进度
- 在模拟测试中比其他方法更有效完成治疗目标且对话更高效
- 适合研究智能心理辅导系统或想提升对话结构化的开发者
基于大语言模型(LLM)的咨询代理虽能生成流畅回应,但缺乏结构化、目标导向的进展。我们提出DeepSAGE(战略人工智能引导引擎),一个结合LLM与深度强化学习(DRL)的混合框架,基于认知行为疗法(CBT)第一阶段设计了11个明确治疗目标的对话阶段。外部控制器判断阶段完成情况,DRL模型选择指导响应生成的治疗意图。在六种检索、提示、阶段和策略基线对比中,DeepSAGE在模拟客户参与度和开放性上表现更好,且在阶段目标达成与对话效率之间取得最佳平衡。领域专家评审认为生成对话具有合理的情绪发展轨迹和可识别的CBT过程。由于评估主要依赖模拟客户和模型指标,结果表明是对话控制能力的相对提升,而非临床有效性。这表明结合阶段结构与学习策略选择是构建AI心理咨询的有前景方向,但其临床效果、安全性和实际应用仍需进一步人工评估。
原文摘要 · Abstract (English)
Large Language Model (LLM)-based counseling agents can generate fluent and supportive responses, but they often lack the structured, goal-directed progression required to conduct a coherent therapeutic session. We present DeepSAGE (Strategic AI Guidance Engine), a hybrid LLM--Deep Reinforcement Learning (DRL) framework for stage-aware counseling dialogue grounded in the first session of Cognitive Behavioral Therapy (CBT). DeepSAGE represents the session as eleven stages with explicit therapeutic objectives, with an external controller determines stage completion and the DRL model selects therapeutic intentions that guide LLM response generation. We evaluate DeepSAGE against six retrieval-, prompting-, stage-, and policy-based alternatives. DeepSAGE elicits higher simulated client engagement and openness and achieves the strongest balance of stage-goal completion and dialogue efficiency among stage-structured systems. Domain expert review further indicates that the generated conversations exhibit broadly plausible emotional trajectories and recognizable CBT processes. Because the evaluation relies primarily on simulated clients and model-based metrics, these findings demonstrate comparative dialogue-control improvements rather than clinical effectiveness. These results suggest that combining stage-structured dialogue with learned strategy selection is a promising approach for AI counseling, though clinical effectiveness, safety, and real-world utility require further human evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。