arXiv:2504.14657cs.CLcs.AI2025-04中稿 · the Conference of …被引 15

用商业大模型生成医疗记录,小规模可行但高维数据难保真实分布。

A Case Study Exploring the Current Landscape of Synthetic Medical Record Generation with Commercial LLMs

  • 用商用大模型生成医疗记录,小特征集可稳定输出。
  • 特征维度增加时,数据分布与相关性严重失真。
  • 适合测试小规模数据生成,不适用于跨院通用场景。

合成电子健康记录(EHR)为创建隐私保护且结构统一的数据提供了宝贵机会,支持医疗领域的多种应用。合成数据的优势包括对数据模式的精确控制、提升患者群体的公平性与代表性,以及在不泄露真实个体隐私的前提下共享数据。因此,人工智能社区越来越多地使用大型语言模型(LLMs)在多个领域生成合成数据。然而,医疗领域一个关键挑战是确保合成记录在不同医院间具有可靠的泛化能力,这一直是该领域的长期难题。本文评估了当前商用大模型生成合成数据的能力,并研究了生成过程中的多个方面,以识别其优势与不足。主要发现表明:尽管大模型能可靠生成小规模特征子集的合成记录,但在数据维度增加时,难以保持真实的分布和相关性,最终限制了其在多样化医院环境中的泛化能力。

原文摘要 · Abstract (English)

Synthetic Electronic Health Records (EHRs) offer a valuable opportunity to create privacy preserving and harmonized structured data, supporting numerous applications in healthcare. Key benefits of synthetic data include precise control over the data schema, improved fairness and representation of patient populations, and the ability to share datasets without concerns about compromising real individuals privacy. Consequently, the AI community has increasingly turned to Large Language Models (LLMs) to generate synthetic data across various domains. However, a significant challenge in healthcare is ensuring that synthetic health records reliably generalize across different hospitals, a long standing issue in the field. In this work, we evaluate the current state of commercial LLMs for generating synthetic data and investigate multiple aspects of the generation process to identify areas where these models excel and where they fall short. Our main finding from this work is that while LLMs can reliably generate synthetic health records for smaller subsets of features, they struggle to preserve realistic distributions and correlations as the dimensionality of the data increases, ultimately limiting their ability to generalize across diverse hospital settings.

合成数据医疗生成大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。