用10类治疗动作分析LLM如何做心理咨询,发现模型爱提问、缺教育、依赖人类引导。
Move by Move: Measuring and Steering How LLMs Conduct Psychotherapy
- 构建10类心理治疗动作分类体系,经专家标注验证并可扩展。
- 模型提问频率是人类3倍,心理教育内容严重不足,策略多跟从人类。
- 用该框架指导模型后,与人类行为差距减半,对话对齐度提升7-9%。
用户越来越多地向大语言模型寻求情感支持,但对其如何开展心理治疗对话仍知之甚少。本文提出一个包含十类治疗动作的本体论:基于MULTI-60量表的紧凑功能类别,经五位持证心理医生标注验证,并通过评分员方法实现专家一致性对齐。将该本体应用于真实咨询记录和模型主导会话,对比人类临床医生与前沿模型的动作用分布。结果显示,模型提问频率最高达人类三倍,却显著忽略心理教育,且高度依赖上下文:能延续人类发起的策略,但极少主动发起。将该本体作为工具引入后,模型平均偏离人类动作分布的幅度减半,回合级对齐度提升7-9个百分点,无需任何微调。
原文摘要 · Abstract (English)
Users increasingly turn to large language models for emotional support, yet little is known about how these models actually conduct a psychotherapy interaction. We introduce an ontology of ten therapeutic moves: compact, function-based categories grounded in the MULTI-60 inventory, validated through an annotation campaign with five licensed psychologists, and scaled with a judge-based approach that matches expert agreement. Applying it to real counseling transcripts and model-led sessions, we compare the move distributions between human clinicians and a panel of frontier models. Models over-use inquiry at up to three times the human rate, neglect psychoeducation, and are strongly context-anchored: they carry forward strategies initiated by a human clinician but rarely initiate them themselves. Exposing the ontology as a set of tools roughly halves the mean deviation from the human move distribution and improves turn-level alignment with human therapist by 7-9 percentage points, without any fine-tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。