arXiv:2503.13509cs.LGcs.AI2025-03KDD被引 58

构建16K对话数据集,推动心理健康对话AI研究

MentalChat16K: A Benchmark Dataset for Conversational Mental Health Assistance

  • 融合合成数据与真实护理对话,覆盖抑郁焦虑等常见心理问题
  • 包含16,000条匿名对话,支持大模型在心理支持场景的训练与评估
  • 注重隐私保护,适合研究共情型AI与医疗心理健康应用

我们提出MentalChat16K,一个英语基准数据集,整合了合成的心理健康咨询数据与真实护工与临终关怀患者照护者之间的匿名对话记录。涵盖抑郁、焦虑、悲伤等多种心理状况,该数据集旨在推动对话式心理健康辅助大语言模型的研发与评估。通过提供高质量、针对性强的资源,MentalChat16K致力于促进更具同理心、个性化的AI解决方案发展,以提升心理支持服务的可及性。数据集严格遵循患者隐私保护、伦理规范和负责任的数据使用原则。该数据集已公开于Hugging Face(https://huggingface.co/datasets/ShenLab/MentalChat16K),代码与文档托管于GitHub(https://github.com/ChiaPatricia/MentalChat16K)。

原文摘要 · Abstract (English)

We introduce MentalChat16K, an English benchmark dataset combining a synthetic mental health counseling dataset and a dataset of anonymized transcripts from interventions between Behavioral Health Coaches and Caregivers of patients in palliative or hospice care. Covering a diverse range of conditions like depression, anxiety, and grief, this curated dataset is designed to facilitate the development and evaluation of large language models for conversational mental health assistance. By providing a high-quality resource tailored to this critical domain, MentalChat16K aims to advance research on empathetic, personalized AI solutions to improve access to mental health support services. The dataset prioritizes patient privacy, ethical considerations, and responsible data usage. MentalChat16K presents a valuable opportunity for the research community to innovate AI technologies that can positively impact mental well-being. The dataset is available at https://huggingface.co/datasets/ShenLab/MentalChat16K and the code and documentation are hosted on GitHub at https://github.com/ChiaPatricia/MentalChat16K.

心理健康对话数据大模型隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。