arXiv:2412.05153cs.LG2024-12中稿 · the 2025 IEEE Inte…被引 13

用大模型凭描述生成真实患者数据,无需原始数据或机器学习基础。

A text-to-tabular approach to generate synthetic patient data using LLMs

  • 仅需数据库描述,利用大模型上下文学习生成数据
  • 生成数据保留临床关联性,隐私与真实性表现良好
  • 适合快速定制数据、科研验证和医学教育使用

获取大规模高质量医疗数据库是加速医学研究和发现疾病规律的关键,但常受限于患者隐私、数据共享限制和高昂成本。为此,合成患者数据成为替代方案。然而,传统合成数据生成(SDG)方法依赖在原始数据上训练的机器学习模型,仍面临数据稀缺问题。本文提出一种无需原始数据、仅需目标数据库描述即可生成合成表格型患者数据的方法。利用大语言模型(LLMs)的先验医学知识与上下文学习能力,在低资源环境下生成具有真实性的患者数据。通过保真度、隐私性和实用性指标评估,结果表明:尽管性能不及基于原始数据训练的顶尖模型,但所生成数据能有效保留临床关联性。消融实验揭示提示设计中关键要素对生成质量的影响。该方法操作简便,无需原始数据或高级机器学习技能,特别适用于快速生成定制化数据,支持项目落地与教学应用。

原文摘要 · Abstract (English)

Access to large-scale high-quality healthcare databases is key to accelerate medical research and make insightful discoveries about diseases. However, access to such data is often limited by patient privacy concerns, data sharing restrictions and high costs. To overcome these limitations, synthetic patient data has emerged as an alternative. However, synthetic data generation (SDG) methods typically rely on machine learning (ML) models trained on original data, leading back to the data scarcity problem. We propose an approach to generate synthetic tabular patient data that does not require access to the original data, but only a description of the desired database. We leverage prior medical knowledge and in-context learning capabilities of large language models (LLMs) to generate realistic patient data, even in a low-resource setting. We quantitatively evaluate our approach against state-of-the-art SDG models, using fidelity, privacy, and utility metrics. Our results show that while LLMs may not match the performance of state-of-the-art models trained on the original data, they effectively generate realistic patient data with well-preserved clinical correlations. An ablation study highlights key elements of our prompt contributing to high-quality synthetic patient data generation. This approach, which is easy to use and does not require original data or advanced ML skills, is particularly valuable for quickly generating custom-designed patient data, supporting project implementation and providing educational resources.

合成数据大模型医疗数据隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。