用大模型模拟专业心理咨询,提升健康行为干预可及性
Toward expert-level motivational interviewing for health behavior improvement with LLMs
- 用中文心理对话数据微调大模型,生成符合动机访谈规范的对话
- 微调后模型在技术与关系维度得分接近真人咨询水平
- 适合需要低成本、可扩展心理支持的医疗与健康管理场景
动机访谈(MI)是一种有效的健康行为干预方法,但受限于对专业咨询师的依赖。本研究探索通过大语言模型实现可扩展的替代方案,构建并评估了面向动机访谈的中文大模型(MI-LLMs)。研究首先整合五个中文心理辅导语料库,利用GPT-4基于动机访谈指导提示,从两个高质量数据集(CPsyCounD和PsyDTCorpus)中生成2,040段多轮动机访谈式对话,其中2,000段用于训练,40段用于测试。选用三个中文开源模型(Baichuan2-7B-Chat、ChatGLM-4-9B-Chat、Llama-3-8B-Chinese-Chat-v2)在该语料上进行微调,命名为MI-LLMs。通过轮次级自动指标和专家人工编码(采用MITI 4.2.1编码手册)评估模型表现。结果表明,微调显著提升了各模型的BLEU-4与ROUGE分数;人工评估显示,MI-LLMs在技术与关系维度的总分及符合动机访谈比例接近真实对话水平,但复杂反思与反思转提问频率仍较低。研究提供初步证据表明,基于动机访谈目标的微调可使通用大模型具备核心符合动机访谈的行为特征,为人工智能辅助健康行为改变提供了可扩展路径,但仍需在数据规模、复杂技能建模及真实干预试验方面进一步研究。
原文摘要 · Abstract (English)
Background: Motivational interviewing (MI) is an effective counseling approach for promoting health behavior change, but its impact is constrained by the need for highly trained human counselors. Objective: This study aimed to explore a scalable alternative by developing and evaluating Large Language Models for Motivational Interviewing (MI-LLMs). Methods: We first curated five Chinese psychological counseling corpora and, using GPT-4 with an MI-informed prompt, transcribed multi-turn dialogues from the two highest-quality datasets (CPsyCounD and PsyDTCorpus) into 2,040 MI-style counseling conversations, of which 2,000 were used for training and 40 for testing. Three Chinese-capable open-source LLMs (Baichuan2-7B-Chat, ChatGLM-4-9B-Chat and Llama-3-8B-Chinese-Chat-v2) were fine-tuned on this corpus and were named as MI-LLMs. We evaluated MI-LLMs using round-based automatic metrics and expert manual coding with the Motivational Interviewing Treatment Integrity (MITI) Coding Manual 4.2.1. Results: Across all three models, fine-tuning substantially improved BLEU-4 and ROUGE scores compared with the base models, and manual coding showed that MI-LLMs achieved technical and relational global scores, and MI-adherent ratios that approached those of real MI dialogues, although complex reflections and reflection-to-question ratios remained less frequent. Conclusions: These findings provide initial evidence that MI-oriented fine-tuning can endow general-purpose LLMs with core MI-consistent counseling behaviors, suggesting a scalable pathway toward AI-assisted health behavior change support while underscoring the need for further work on data scale, complex MI skills and real-world intervention trials.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。