构建28种情绪标签的日本医疗共情对话数据集,助力精准情感响应。
EmplifAI: a Fine-grained Dataset for Japanese Empathetic Medical Dialogues in 28 Emotion Labels
- 基于28类细粒度情绪标签构建医疗情境对话
- 4125条对话在多模型上实现0.83的BERTScore F1
- 可评估模型共情能力,适合医疗AI研究者
本文提出EmplifAI,一个面向慢性病患者心理支持的日本共情对话数据集。针对患者在疾病管理不同阶段经历的丰富正负情绪(如希望与绝望),数据集涵盖280个医学情境和4125组两轮对话,基于GoEmotions分类体系细化为28种情绪标签,通过众包与专家评审收集。为评估对话情感契合度,使用BERTScore对多个大语言模型(LLMs)在情境-对话对上的预测进行评估,获得0.83的F1分数。以基线日语模型LLM-jp-3.1-13b-instruct4在EmplifAI上微调后,显著提升流畅性、通用共情及情绪特异性共情表现。同时,对比了LLM-as-a-Judge与人工评分在多模型生成对话上的评分结果,验证评估流程有效性,并分析相关性带来的启示与潜在风险。
原文摘要 · Abstract (English)
This paper introduces EmplifAI, a Japanese empathetic dialogue dataset designed to support patients coping with chronic medical conditions. They often experience a wide range of positive and negative emotions (e.g., hope and despair) that shift across different stages of disease management. EmplifAI addresses this complexity by providing situation-based dialogues grounded in 28 fine-grained emotion categories, adapted and validated from the GoEmotions taxonomy. The dataset includes 280 medically contextualized situations and 4125 two-turn dialogues, collected through crowdsourcing and expert review. To evaluate emotional alignment in empathetic dialogues, we assessed model predictions on situation--dialogue pairs using BERTScore across multiple large language models (LLMs), achieving F1 scores of 0.83. Fine-tuning a baseline Japanese LLM (LLM-jp-3.1-13b-instruct4) with EmplifAI resulted in notable improvements in fluency, general empathy, and emotion-specific empathy. Furthermore, we compared the scores assigned by LLM-as-a-Judge and human raters on dialogues generated by multiple LLMs to validate our evaluation pipeline and discuss the insights and potential risks derived from the correlation analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。