首份系统梳理临床心理健康AI数据集的综述,助力模型更可靠、公平。
A Comprehensive Review of Datasets for Clinical Mental Health AI Systems
- 按疾病类型、数据模态、任务等维度分类整理现有数据集
- 发现长期追踪数据少、文化语言代表性不足等关键缺口
- 适合研究者构建更普适、可复现的心理健康AI系统
全球精神健康问题日益严重,但专业临床人员供给未能同步增长,导致大量人群无法获得及时支持。人工智能(AI)在精神健康诊断、监测与干预中展现出潜力,但其有效性高度依赖高质量临床训练数据。然而,当前相关数据集仍分散、文档不全且难获取,严重影响模型的可复现性、可比性和泛化能力。本文首次系统综述了用于训练临床心理AI助手的各类数据集,按精神障碍类型(如抑郁症、精神分裂症)、数据模态(如文本、语音、生理信号)、任务类型(如诊断预测、症状严重度评估、干预生成)、可访问性(公开、受限或私有)及社会文化背景(如语言与文化)进行分类。同时探讨了合成临床数据集。研究揭示了长期数据缺失、文化语言代表性不足、采集与标注标准不一、合成数据模态有限等关键问题。最后提出未来数据集构建与标准化的关键挑战,并给出具体建议,以推动更稳健、泛化能力强且公平的心理健康AI系统发展。
原文摘要 · Abstract (English)
Mental health disorders are rising worldwide. However, the availability of trained clinicians has not scaled proportionally, leaving many people without adequate or timely support. To bridge this gap, recent studies have shown the promise of Artificial Intelligence (AI) to assist mental health diagnosis, monitoring, and intervention. However, the development of efficient, reliable, and ethical AI to assist clinicians is heavily dependent on high-quality clinical training datasets. Despite growing interest in data curation for training clinical AI assistants, existing datasets largely remain scattered, under-documented, and often inaccessible, hindering the reproducibility, comparability, and generalizability of AI models developed for clinical mental health care. In this paper, we present the first comprehensive survey of clinical mental health datasets relevant to the training and development of AI-powered clinical assistants. We categorize these datasets by mental disorders (e.g., depression, schizophrenia), data modalities (e.g., text, speech, physiological signals), task types (e.g., diagnosis prediction, symptom severity estimation, intervention generation), accessibility (public, restricted or private), and sociocultural context (e.g., language and cultural background). Along with these, we also investigate synthetic clinical mental health datasets. Our survey identifies critical gaps such as a lack of longitudinal data, limited cultural and linguistic representation, inconsistent collection and annotation standards, and a lack of modalities in synthetic data. We conclude by outlining key challenges in curating and standardizing future datasets and provide actionable recommendations to facilitate the development of more robust, generalizable, and equitable mental health AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。