用行为指标评估大模型心理咨询能力,效果接近专业治疗师。
AI-Augmented LLMs Achieve Therapist-Level Responses in Motivational Interviewing
- 构建17项行为指标,区分符合与不符合咨询规范的对话
- 优化提示词后,大模型在共情和反思上显著提升
- 适合临床辅助工具研发者及人机协作研究者参考
大型语言模型(如GPT-4)在成瘾干预中具有规模化动机访谈(MI)的潜力,但需系统评估其治疗能力。本文提出一种计算框架,通过预期与非预期的MI行为评估用户感知质量(UPQ)。基于人类治疗师与GPT-4的对话分析,结合深度学习与可解释AI,建立预测模型,识别出17个符合(MICO)与不符合(MIIN)MI规范的行为指标。定制化的链式思维提示显著提升了GPT-4的表现,减少不当建议,增强共情与反思。尽管整体仍略逊于人类治疗师,但在建议管理方面表现更优。该框架通过行为指标分析与人机协同评估,为优化基于LLM的治疗工具提供了路径。研究揭示了大模型在临床沟通应用中的可扩展性与当前局限。
原文摘要 · Abstract (English)
Large language models (LLMs) like GPT-4 show potential for scaling motivational interviewing (MI) in addiction care, but require systematic evaluation of therapeutic capabilities. We present a computational framework assessing user-perceived quality (UPQ) through expected and unexpected MI behaviors. Analyzing human therapist and GPT-4 MI sessions via human-AI collaboration, we developed predictive models integrating deep learning and explainable AI to identify 17 MI-consistent (MICO) and MI-inconsistent (MIIN) behavioral metrics. A customized chain-of-thought prompt improved GPT-4's MI performance, reducing inappropriate advice while enhancing reflections and empathy. Although GPT-4 remained marginally inferior to therapists overall, it demonstrated superior advice management capabilities. The model achieved measurable quality improvements through prompt engineering, yet showed limitations in addressing complex emotional nuances. This framework establishes a pathway for optimizing LLM-based therapeutic tools through targeted behavioral metric analysis and human-AI co-evaluation. Findings highlight both the scalability potential and current constraints of LLMs in clinical communication applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。