为心理治疗笔记质量评估建立新标准,发现AI生成笔记反而更受医生青睐。
TN-Eval: Rubric and Evaluation Protocols for Measuring the Quality of Behavioral Therapy Notes
- 设计涵盖完整性、简洁性、忠实性的评估量规,替代传统评分方式。
- 大模型在完整性和简洁性上可媲真人,但存在幻觉问题。
- 医生盲测更偏好AI生成笔记,提示其实际应用潜力。
心理治疗笔记对法律合规和患者照护至关重要,但其质量标准尚不完善。本文与持证治疗师合作,构建涵盖完整性、简洁性和忠实性的综合评估量规,并扩展公开的心理健康对话数据集,包含治疗师撰写和大模型生成的笔记。应用该框架评估后发现:(1)基于量规的手动评估比传统李克特量表更可靠、可解释;(2)大模型能有效模拟人类对完整性和简洁性的判断,但在忠实性上表现不佳;(3)治疗师笔记普遍存在不完整和冗长问题,而大模型生成笔记存在幻觉现象。令人意外的是,在盲测中,治疗师更偏好并认为大模型生成笔记质量更高。
原文摘要 · Abstract (English)
Behavioral therapy notes are important for both legal compliance and patient care. Unlike progress notes in physical health, quality standards for behavioral therapy notes remain underdeveloped. To address this gap, we collaborated with licensed therapists to design a comprehensive rubric for evaluating therapy notes across key dimensions: completeness, conciseness, and faithfulness. Further, we extend a public dataset of behavioral health conversations with therapist-written notes and LLM-generated notes, and apply our evaluation framework to measure their quality. We find that: (1) A rubric-based manual evaluation protocol offers more reliable and interpretable results than traditional Likert-scale annotations. (2) LLMs can mimic human evaluators in assessing completeness and conciseness but struggle with faithfulness. (3) Therapist-written notes often lack completeness and conciseness, while LLM-generated notes contain hallucination. Surprisingly, in a blind test, therapists prefer and judge LLM-generated notes to be superior to therapist-written notes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。