arXiv:2508.08529cs.AI2025-08被引 5

用提示工程让大模型生成安全可靠的医疗表格数据。

SynLLM: A Comparative Analysis of Large Language Models for Medical Tabular Synthetic Data Generation via Prompt Engineering

  • 设计四类提示模板,不微调模型即可控制生成内容。
  • 规则类提示在隐私与质量间取得最佳平衡。
  • 适合医疗数据研究者和隐私保护需求高的场景。

由于隐私法规限制,真实医疗数据获取困难,制约了医疗研究发展。合成数据提供了替代方案,但生成既真实又符合临床逻辑且保护隐私的记录仍具挑战。本文提出SynLLM框架,利用20个开源大模型(包括LLaMA、Mistral和GPT系列)结合结构化提示生成高质量医疗表格数据。设计四种提示类型:从示例驱动到基于规则的约束,嵌入数据模式、元数据与领域知识以控制输出,无需模型微调。构建全面评估流程,从统计保真度、临床一致性到隐私保护多维度验证。在Diabetes、Cirrhosis、Stroke三个公开数据集上测试,结果表明提示工程显著影响数据质量与隐私风险,规则型提示实现最优平衡。研究表明,经合理提示设计与多指标评估,大模型可生成兼具临床合理性与隐私保护性的合成数据,推动医疗研究更安全高效的数据共享。

原文摘要 · Abstract (English)

Access to real-world medical data is often restricted due to privacy regulations, posing a significant barrier to the advancement of healthcare research. Synthetic data offers a promising alternative; however, generating realistic, clinically valid, and privacy-conscious records remains a major challenge. Recent advancements in Large Language Models (LLMs) offer new opportunities for structured data generation; however, existing approaches frequently lack systematic prompting strategies and comprehensive, multi-dimensional evaluation frameworks. In this paper, we present SynLLM, a modular framework for generating high-quality synthetic medical tabular data using 20 state-of-the-art open-source LLMs, including LLaMA, Mistral, and GPT variants, guided by structured prompts. We propose four distinct prompt types, ranging from example-driven to rule-based constraints, that encode schema, metadata, and domain knowledge to control generation without model fine-tuning. Our framework features a comprehensive evaluation pipeline that rigorously assesses generated data across statistical fidelity, clinical consistency, and privacy preservation. We evaluate SynLLM across three public medical datasets, including Diabetes, Cirrhosis, and Stroke, using 20 open-source LLMs. Our results show that prompt engineering significantly impacts data quality and privacy risk, with rule-based prompts achieving the best privacy-quality balance. SynLLM establishes that, when guided by well-designed prompts and evaluated with robust, multi-metric criteria, LLMs can generate synthetic medical data that is both clinically plausible and privacy-aware, paving the way for safer and more effective data sharing in healthcare research.

医疗数据大模型合成数据提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。