构建首个韩语动机访谈对话数据集,助力心理聊天机器人发展
KMI: A Dataset of Korean Motivational Interviewing Dialogues for Psychotherapy
- 用专业治疗师行为模拟训练模型,结合LLM生成对话
- 产出1000条高质量韩语动机访谈对话,经专家评估验证
- 适合心理AI、跨语言对话系统研究者使用
心理健康服务需求上升推动了AI心理聊天机器人的发展,但隐私、数据收集和专业性仍是挑战。动机访谈(MI)因其理论基础日益被用于提升聊天机器人专业能力。然而现有数据集存在局限,尤其在非英语语言中更为匮乏。本文提出一种融合专业治疗师经验的模拟框架,训练一个模仿治疗师行为选择的MI预测模型,并利用大语言模型通过提示工程生成对话。最终构建了首个基于MI理论的合成数据集KMI,包含1000条高质量韩语动机访谈对话。通过专家评估及在该数据集上训练的对话模型测试,验证了KMI在质量、专业性和实用性上的优势。同时引入源自MI理论的新评价指标,从理论角度评估对话效果。
原文摘要 · Abstract (English)
The increasing demand for mental health services has led to the rise of AI-driven mental health chatbots, though challenges related to privacy, data collection, and expertise persist. Motivational Interviewing (MI) is gaining attention as a theoretical basis for boosting expertise in the development of these chatbots. However, existing datasets are showing limitations for training chatbots, leading to a substantial demand for publicly available resources in the field of MI and psychotherapy. These challenges are even more pronounced in non-English languages, where they receive less attention. In this paper, we propose a novel framework that simulates MI sessions enriched with the expertise of professional therapists. We train an MI forecaster model that mimics the behavioral choices of professional therapists and employ Large Language Models (LLMs) to generate utterances through prompt engineering. Then, we present KMI, the first synthetic dataset theoretically grounded in MI, containing 1,000 high-quality Korean Motivational Interviewing dialogues. Through an extensive expert evaluation of the generated dataset and the dialogue model trained on it, we demonstrate the quality, expertise, and practicality of KMI. We also introduce novel metrics derived from MI theory in order to evaluate dialogues from the perspective of MI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。