arXiv:2411.17672cs.LG2024-11被引 27

用大模型生成抑郁检测合成数据,解决隐私与数据不足问题

Synthetic Data Generation with LLM for Improved Depression Prediction

  • 用链式思维提示生成临床访谈文本的摘要与情感分析
  • 合成数据有效平衡严重程度分布,提升模型预测能力
  • 适合心理健康、隐私保护与小样本学习研究者

自动抑郁检测是心理学与机器学习交叉的热门领域,但因话题敏感导致数据隐私与稀缺问题日益突出。本文提出一种基于大语言模型(LLMs)的合成数据生成流程,从临床访谈的非结构化文本出发,通过链式思维提示生成摘要与情感分析。流程分两步:第一步基于原始转录文本和抑郁评分生成摘要与情感分析;第二步基于第一步结果与新抑郁评分生成合成摘要与情感分析。合成数据在保真度与隐私保护指标上表现良好,同时平衡了训练集中的症状严重程度分布,显著提升了模型对患者抑郁强度的预测能力。该方法通过扩充真实世界中有限且不平衡的数据集,为自动抑郁检测中的数据稀缺与隐私问题提供了新思路,且保持了原始数据的统计完整性,可为未来心理健康研究提供稳健框架。

原文摘要 · Abstract (English)

Automatic detection of depression is a rapidly growing field of research at the intersection of psychology and machine learning. However, with its exponential interest comes a growing concern for data privacy and scarcity due to the sensitivity of such a topic. In this paper, we propose a pipeline for Large Language Models (LLMs) to generate synthetic data to improve the performance of depression prediction models. Starting from unstructured, naturalistic text data from recorded transcripts of clinical interviews, we utilize an open-source LLM to generate synthetic data through chain-of-thought prompting. This pipeline involves two key steps: the first step is the generation of the synopsis and sentiment analysis based on the original transcript and depression score, while the second is the generation of the synthetic synopsis/sentiment analysis based on the summaries generated in the first step and a new depression score. Not only was the synthetic data satisfactory in terms of fidelity and privacy-preserving metrics, it also balanced the distribution of severity in the training dataset, thereby significantly enhancing the model's capability in predicting the intensity of the patient's depression. By leveraging LLMs to generate synthetic data that can be augmented to limited and imbalanced real-world datasets, we demonstrate a novel approach to addressing data scarcity and privacy concerns commonly faced in automatic depression detection, all while maintaining the statistical integrity of the original dataset. This approach offers a robust framework for future mental health research and applications.

抑郁检测合成数据大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。