为心理治疗聊天机器人设计了可量化的临床真实与安全评估框架
TherapyGym: Evaluating and Aligning Clinical Fidelity and Safety in Therapy Chatbots
- 用自动化工具按认知行为疗法标准评分对话质量
- 模型经训练后临床评分从0.10提升至0.60,显著改善
- 适合开发需符合专业诊疗规范的心理健康AI
大型语言模型在心理健康支持中应用日益广泛,但现有评价方法——流畅性指标、偏好测试和通用对话基准——难以捕捉心理治疗中的临床关键维度。本文提出THERAPYGYM框架,从临床真实性和安全性双维度评估治疗类聊天机器人。真实性采用认知行为疗法评分量表(CTRS)进行自动化多轮对话评估;安全性则通过多标签标注方案覆盖治疗特定风险(如忽视伤害或虐待)。为减少大模型评判的偏差与不可靠性,还发布了THERAPYJUDGEBENCH,包含116段对话及1,270条专家评级数据,用于审计与校准。THERAPYGYM同时作为训练工具:基于CTRS与安全性的奖励信号,结合多种症状特征的患者模拟环境进行强化学习。在该框架下训练的模型在专家评分中表现提升,平均CTRS从0.10升至0.60(大模型裁判下从0.16升至0.59),验证了其对循证实践的忠实性与高风险场景下的安全性提升能力。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used for mental-health support; yet prevailing evaluation methods--fluency metrics, preference tests, and generic dialogue benchmarks--fail to capture the clinically critical dimensions of psychotherapy. We introduce THERAPYGYM, a framework that evaluates and improves therapy chatbots along two clinical pillars: fidelity and safety. Fidelity is measured using the Cognitive Therapy Rating Scale (CTRS), implemented as an automated pipeline that scores adherence to CBT techniques over multi-turn sessions. Safety is assessed using a multi-label annotation scheme, covering therapy-specific risks (e.g., failing to address harm or abuse). To mitigate bias and unreliability in LLM-based judges, we further release THERAPYJUDGEBENCH, a validation set of 116 dialogues with 1,270 expert ratings for auditing and calibration against licensed clinicians. THERAPYGYM also serves as a training harness: CTRS and safety-based rewards drive RL with configurable patient simulations spanning diverse symptom profiles. Models trained in THERAPYGYM improve on expert ratings, with average CTRS rising from 0.10 to 0.60 (and 0.16 to 0.59 under LLM judges). Our work enables scalable development of therapy chatbots that are faithful to evidence-based practice and safer in high-stakes use.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。