arXiv:2604.05795cs.CL2026-04ACL

提出新评估框架CARE,精准衡量AI对话是否符合心理治疗核心原则

Measuring What Matters!! Assessing Therapeutic Principles in Mental-Health Conversation

  • 构建六维临床准则评分体系,结合专家标注的有序等级
  • 新框架CARE在测试中F1达63.34,较基线提升64.26%
  • 适合关注AI心理服务可信度的研究者与开发者

大型语言模型在心理健康应用中的普及,亟需超越表面流畅性的系统性评估框架,以确保其符合心理治疗最佳实践。本文研究AI生成治疗型对话在临床合理性与有效性方面的评估问题,从非评判接纳、温暖、尊重自主性、积极倾听、反映理解、情境恰当性六个核心治疗原则出发,采用细粒度有序量表进行评估。我们引入FAITH-M基准,包含专家赋予的有序评分;并提出CARE多阶段评估框架,融合对话内上下文建模、对比示例检索与知识蒸馏的思维链推理。实验表明,CARE取得63.34的F1分数,显著优于基线模型Qwen3的38.56(提升64.26%),且优势来自结构化推理与上下文建模而非模型规模本身。专家评估与外部数据集验证显示其在领域迁移下的鲁棒性,但仍面临隐含临床细节建模的挑战。总体而言,CARE为评估AI心理健康系统中的治疗契合度提供了临床基础框架。

原文摘要 · Abstract (English)

The increasing use of large language models in mental health applications calls for principled evaluation frameworks that assess alignment with psychotherapeutic best practices beyond surface-level fluency. While recent systems exhibit conversational competence, they lack structured mechanisms to evaluate adherence to core therapeutic principles. In this paper, we study the problem of evaluating AI-generated therapist-like responses for clinically grounded appropriateness and effectiveness. We assess each therapists utterance along six therapeutic principles: non-judgmental acceptance, warmth, respect for autonomy, active listening, reflective understanding, and situational appropriateness using a fine-grained ordinal scale. We introduce FAITH-M, a benchmark annotated with expert-assigned ordinal ratings, and propose CARE, a multi-stage evaluation framework that integrates intra-dialogue context, contrastive exemplar retrieval, and knowledge-distilled chain-of-thought reasoning. Experiments show that CARE achieves an F-1 score of 63.34 versus the strong baseline Qwen3 F-1 score of 38.56 which is a 64.26 improvement, which also serves as its backbone, indicating that gains arise from structured reasoning and contextual modeling rather than backbone capacity alone. Expert assessment and external dataset evaluations further demonstrate robustness under domain shift, while highlighting challenges in modelling implicit clinical nuance. Overall, CARE provides a clinically grounded framework for evaluating therapeutic fidelity in AI mental health systems.

心理AI评估框架对话系统临床评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。