用临床指南引导大模型生成隐私安全的精神病数据,无需真实数据即可保持高质量。
Knowledge-Guided Retrieval-Augmented Generation for Zero-Shot Psychiatric Data: Privacy Preserving Synthetic Data Generation
- 用DSM-5和ICD-10知识库引导LLM进行检索增强生成
- 在6种焦虑症中,配对结构误差最低,尤其在社交恐惧和分离焦虑上表现优
- 不依赖真实数据,隐私风险低,适合数据受限场景
医疗AI在提升患者处理效率方面展现出潜力,但受限于真实患者数据的获取。为解决此问题,本文提出一种零样本、知识引导的精神病表格数据生成框架,利用大语言模型(LLMs)结合《精神疾病诊断与统计手册》(DSM-5)和《国际疾病分类》(ICD-10)进行检索增强生成。通过不同知识库组合生成隐私保护的合成数据,并与两种依赖真实数据的先进深度学习模型(CTGAN和TVAE)对比。评估涵盖六种焦虑相关障碍:特定恐惧症、社交焦虑障碍、广场恐惧症、广泛性焦虑障碍、分离焦虑障碍和恐慌障碍。结果表明,虽然CTGAN在边缘分布和多变量结构上表现最佳,但知识增强型LLM在配对结构上具有竞争力,且在分离焦虑和社交焦虑中达到最低的配对误差。消融实验显示,临床检索显著提升单变量与配对保真度。隐私分析表明,该无真实数据的LLM模型重叠度较低,平均链接风险与CTGAN相当;而TVAE虽有低k-map得分,却存在大量重复。总体而言,在缺乏真实数据或无法共享时,将LLM基于临床知识可生成高质量且隐私安全的合成精神病数据。
原文摘要 · Abstract (English)
AI systems in healthcare research have shown potential to increase patient throughput and assist clinicians, yet progress is constrained by limited access to real patient data. To address this issue, we present a zero-shot, knowledge-guided framework for psychiatric tabular data in which large language models (LLMs) are steered via Retrieval-Augmented Generation using the Diagnostic and Statistical Manual of Mental Disorders (DSM-5) and the International Classification of Diseases (ICD-10). We conducted experiments using different combinations of knowledge bases to generate privacy-preserving synthetic data. The resulting models were benchmarked against two state-of-the-art deep learning models for synthetic tabular data generation, namely CTGAN and TVAE, both of which rely on real data and therefore entail potential privacy risks. Evaluation was performed on six anxiety-related disorders: specific phobia, social anxiety disorder, agoraphobia, generalized anxiety disorder, separation anxiety disorder, and panic disorder. CTGAN typically achieves the best marginals and multivariate structure, while the knowledge-augmented LLM is competitive on pairwise structure and attains the lowest pairwise error in separation anxiety and social anxiety. An ablation study shows that clinical retrieval reliably improves univariate and pairwise fidelity over a no-retrieval LLM. Privacy analyses indicate that the real data-free LLM yields modest overlaps and a low average linkage risk comparable to CTGAN, whereas TVAE exhibits extensive duplication despite a low k-map score. Overall, grounding an LLM in clinical knowledge enables high-quality, privacy-preserving synthetic psychiatric data when real datasets are unavailable or cannot be shared.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。