arXiv:2506.16594cs.CL2025-06综述被引 9

用大模型生成医学数据,解决数据稀缺与质量难题

A Scoping Review of Synthetic Data Generation by Language Models in Biomedical Research and Application: Data Utility and Quality Perspectives

  • 按提示词、微调等方法生成文本、表格等医学数据
  • 78%研究聚焦非结构化文本,44%依赖人工评估质量
  • 适合医学数据不足或需隐私保护的研究者

利用大语言模型(LLMs)生成合成数据在应对生物医学数据挑战方面展现出巨大潜力,并在生物医学研究中得到越来越多应用。本研究遵循PRISMA-ScR指南,系统回顾了2020至2025年间发表于PubMed、ACM、Web of Science和Google Scholar的文献,共纳入59项与生物医学领域合成数据生成相关的研究。其中,主要数据模态为非结构化文本(78.0%)、表格数据(13.6%)和多模态数据(8.4%)。常用生成方法包括提示工程(74.6%)、微调(20.3%)及专用模型(5.1%)。评估方式多样:内在指标占27.1%,人工评估占44.1%,基于LLM的评估占13.6%。然而,数据模态适配性、领域适用性、资源可及性及标准化评估体系仍存在明显障碍。未来应致力于建立透明、统一的评估框架,提升模型与资源的可及性,以支持生物医学研究中的有效应用。

原文摘要 · Abstract (English)

Synthetic data generation using large language models (LLMs) demonstrates substantial promise in addressing biomedical data challenges and shows increasing adoption in biomedical research. This study systematically reviews recent advances in synthetic data generation for biomedical applications and clinical research, focusing on how LLMs address data scarcity, utility, and quality issues with different modalities. We conducted a scoping review following PRISMA-ScR guidelines and searched literature published between 2020 and 2025 through PubMed, ACM, Web of Science, and Google Scholar. A total of 59 studies were included based on relevance to synthetic data generation in biomedical contexts. Among the reviewed studies, the predominant data modalities were unstructured texts (78.0\%), tabular data (13.6\%), and multimodal sources (8.4\%). Common generation methods included LLM prompting (74.6\%), fine-tuning (20.3\%), and specialized models (5.1\%). Evaluations were heterogeneous: intrinsic metrics (27.1\%), human-in-the-loop assessments (44.1\%), and LLM-based evaluations (13.6\%). However, limitations and key barriers persist in data modalities, domain utility, resource and model accessibility, and standardized evaluation protocols. Future efforts may focus on developing standardized, transparent evaluation frameworks and expanding accessibility to support effective applications in biomedical research.

合成数据大模型医学研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。