构建心理访谈对话数据集,提升大模型在心理健康领域的共情生成能力。
Unlocking LLMs: Addressing Scarce Data and Bias Challenges in Mental Health
- 用提示工程生成带情境的对话,结合治疗风格与语义准确性。
- 专家标注数据集,严格遵循心理访谈评分标准,覆盖语言与心理维度。
- 验证大模型在低资源敏感场景下的偏见与情感理解能力,适配临床辅助系统。
大语言模型在医疗分析中展现潜力,但在复杂敏感领域面临幻觉、重复和偏见问题。本文提出IC-AnnoMI,一个基于AnnoMI的专家标注动机访谈(MI)数据集,通过大模型(如ChatGPT)生成包含上下文的对话,利用精心设计的提示策略,结合同理心、反映技巧与语义一致性控制。对话经专家按动机访谈技能编码(MISC)标注,涵盖心理与语言双重维度。我们采用经典机器学习与先进Transformer模型对数据集进行分类任务评估,检验大模型的情感推理与领域理解能力。研究还探讨渐进式提示策略及数据增强对缓解模型偏见的影响。成果为心理治疗领域提供高质量数据集与大模型应用洞察,助力监督环境下共情对话生成。
原文摘要 · Abstract (English)
Large language models (LLMs) have shown promising capabilities in healthcare analysis but face several challenges like hallucinations, parroting, and bias manifestation. These challenges are exacerbated in complex, sensitive, and low-resource domains. Therefore, in this work we introduce IC-AnnoMI, an expert-annotated motivational interviewing (MI) dataset built upon AnnoMI by generating in-context conversational dialogues leveraging LLMs, particularly ChatGPT. IC-AnnoMI employs targeted prompts accurately engineered through cues and tailored information, taking into account therapy style (empathy, reflection), contextual relevance, and false semantic change. Subsequently, the dialogues are annotated by experts, strictly adhering to the Motivational Interviewing Skills Code (MISC), focusing on both the psychological and linguistic dimensions of MI dialogues. We comprehensively evaluate the IC-AnnoMI dataset and ChatGPT's emotional reasoning ability and understanding of domain intricacies by modeling novel classification tasks employing several classical machine learning and current state-of-the-art transformer approaches. Finally, we discuss the effects of progressive prompting strategies and the impact of augmented data in mitigating the biases manifested in IC-AnnoM. Our contributions provide the MI community with not only a comprehensive dataset but also valuable insights for using LLMs in empathetic text generation for conversational therapy in supervised settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。