arXiv:2606.30887cs.CLcs.AI2026-06

用人类对齐的评估驱动治疗型对话生成,提升心理支持质量。

Training Therapeutic Judges and Multi-Agent Systems for Human-Aligned Mental Health Support

论文配图:Training Therapeutic Judges and Multi-Agent Systems for Human-Aligned Mental Health Support
图 1 · 摘自论文原文
  • 构建多维度评估框架,以人类标注数据训练可信赖的治疗评价模型。
  • 引入角色协作系统,将评估反馈转化为精准的响应优化,使评分提升0.43分。
  • 特别擅长修复不安全内容,低质回复恢复率高达94%。

大语言模型在心理支持中展现潜力,但只有当评估作为可操作的控制信号而非被动指标时,治疗质量才能提升。本文提出一个框架,将治疗性回应生成视为由多维度、人类对齐评估驱动的决策优化问题。第一阶段引入TheraJudge,一个基于人类标注数据通过偏好优化训练的开源治疗评估器,在7个心理维度上实现可靠判断,与临床医生评分具有高度一致性(组内相关系数ICC = 0.87–0.95),优于监督基线和强闭源评估器,尤其在安全性、相关性和共情等关键维度表现突出。第二阶段引入TheraAgent,通过具备评论员、教练与治疗师角色的协同修正过程,将评估信号转化为针对性回应修改。实证表明,TheraAgent在盲评下使人类评分提升0.43分(5分制),临床医生评分一致性达96%;低质量回复(≤3分)平均提升2.45分,恢复率达94%,有效纠正不安全输出。结果表明,心理支持大模型的有效对齐依赖于对人类对齐评估的实际响应,而非仅依赖更强的生成能力。代码已开源:https://github.com/vis-nlp/TheraAlign。

原文摘要 · Abstract (English)

Large language models show promise for mental health support, yet therapeutic quality improves only when evaluation functions as an actionable control signal rather than a passive metric. We introduce a framework that formulates therapeutic response generation as a decision-refinement problem driven by multi-dimensional, human-aligned evaluation. In Stage I, we introduce TheraJudge, an open-source therapeutic evaluator trained via preference-based optimization on human-annotated data to produce reliable judgments across 7 psychological dimensions. In Stage II, we introduce TheraAgent, which operationalizes TheraJudge's evaluations through a coordinated refinement process with specialized Critic, Coach, and Therapist roles that translate evaluative signals into targeted response revisions. Empirically, TheraJudge achieves strong agreement with clinician ratings, with intraclass correlation coefficients (ICC = 0.87-0.95), surpassing supervised baselines and strong closed-source judges, particularly on critical dimensions such as Safety, Relevance, and Empathy. Acting on these evaluations, TheraAgent yields a +0.43 improvement in human-rated therapeutic quality (on a 5-point scale) under blind evaluation, with 96\% clinician inter-rater reliability. Low-quality responses ($\leq 3$) improve by +2.45 points with a 94\% recovery rate, demonstrating targeted correction of unsafe outputs. Overall, our results indicate that effective alignment of mental-health LLMs stems from acting on human-aligned evaluation, rather than relying solely on stronger generation. We release code at https://github.com/vis-nlp/TheraAlign.

心理支持评估对齐多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。