用故事化情境提升心理对话生成的临床有效性
StoryMI: Steerable Multi-Agent Therapeutic Dialogue Generation

- 多智能体框架将问卷转为叙事场景,实现对话情境化
- 动态策略控制使12种动机访谈编码符合临床标准
- 适合心理治疗生成与人机交互研究者使用
大型语言模型虽能生成流畅对话,但以往工作缺乏情境锚定、动态策略控制及符合临床标准的评估。本文提出StoryMI,一种基于多语言模型智能体的可控动机访谈对话生成框架。通过问卷生成的故事扩展为情境叙述,为对话提供叙事背景;治疗师与客户智能体在交互智能体调控下,依据选定的动机访谈代码生成对话,后者动态协调交流以控制策略。我们设计双层评估体系:词汇指标与宏观咨询策略的特定度量,辅以大模型评分与人工专家评估。构建包含6000条模拟对话的数据集,基于1000个问卷-故事对,覆盖12种动机访谈编码和13类症状领域,并对六种开源与闭源大模型进行基准测试。结果表明,情境化与宏观策略控制可提升动机访谈依从性与临床合理性,验证了结构化多智能体流程在心理治疗对话生成中的有效性。代码与数据已公开以支持复现。
原文摘要 · Abstract (English)
Large language models (LLMs) can generate fluent dialogue, but prior works lack situational grounding, dynamic strategy control, and evaluation aligned with clinical standards in motivational interviewing (MI). We introduce StoryMI, a multi-LLM agent framework for controllable MI dialogue generation, where questionnaire-based client profiles are expanded into situational stories that provide narrative context for the dialogue. Therapist and client agents generate MI-coded utterances guided by MI codes selected by the interaction agent, while an interaction agent dynamically coordinates exchanges to control MI strategies during a multi-turn conversation. We propose a two-level evaluation protocol: lexical metrics and MI-specific measures of macro-level counseling strategies, alongside LLM-as-judge and human expert assessments. We construct a dataset of 6K simulated MI dialogues grounded in 1K questionnaire-story pairs, covering 12 MI codes and 13 symptom domains, and benchmark six open- and closed-source LLMs. Our results show that situational grounding and macro-level control can improve MI adherence and clinical plausibility, demonstrating the effectiveness of a structured multi-agent workflow for psychotherapy dialogue generation. We provide code and data for reproducibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。