用GPT-4o零样本生成高质量神经外科合成数据,保护隐私且可直接用于模型训练。
Zero-shot generation of synthetic neurosurgical data with large language models
- 直接使用GPT-4o零样本生成神经外科合成数据,无需微调或接触真实数据。
- 生成数据在均值、分布和相关性上高度匹配真实数据,分类任务F1达0.706。
- 适合缺乏真实数据的临床研究者,尤其适用于小样本场景下的机器学习建模。
临床数据对推动神经外科研究至关重要,但受限于数据稀缺、样本量小、隐私法规及耗时的预处理与去标识化流程。合成数据为应对真实世界数据(RWD)获取与使用难题提供了可能。本研究评估了大型语言模型GPT-4o在零样本条件下生成合成神经外科数据的能力,以条件表格式生成对抗网络(CTGAN)为基准进行对比。通过比较合成数据集与真实神经外科数据的保真度(均值、比例、分布及双变量相关性)、实用性(基于RWD的机器学习分类器性能)和隐私性(与RWD记录重复率),发现GPT-4o生成的数据在未进行微调或访问真实数据预训练的情况下,表现优于或媲美CTGAN。生成数据在单变量与双变量保真度上均表现出色,且未暴露任何真实患者记录,即使在扩增样本量后仍保持高一致性。以生成数据训练的分类器在二分类任务中预测术后功能状态恶化,取得F1分数0.706,与基于CTGAN数据训练的结果(0.705)相当。结果表明,GPT-4o具备生成高保真合成神经外科数据的潜力,可有效补充小样本临床数据,并用于预测神经外科结局的机器学习建模。未来需进一步优化分布特性保留并提升分类器性能。
原文摘要 · Abstract (English)
Clinical data is fundamental to advance neurosurgical research, but access is often constrained by data availability, small sample sizes, privacy regulations, and resource-intensive preprocessing and de-identification procedures. Synthetic data offers a potential solution to challenges associated with accessing and using real-world data (RWD). This study aims to evaluate the capability of zero-shot generation of synthetic neurosurgical data with a large language model (LLM), GPT-4o, by benchmarking with the conditional tabular generative adversarial network (CTGAN). Synthetic datasets were compared to real-world neurosurgical data to assess fidelity (means, proportions, distributions, and bivariate correlations), utility (ML classifier performance on RWD), and privacy (duplication of records from RWD). The GPT-4o-generated datasets matched or exceeded CTGAN performance, despite no fine-tuning or access to RWD for pre-training. Datasets demonstrated high univariate and bivariate fidelity to RWD without directly exposing any real patient records, even at amplified sample size. Training an ML classifier on GPT-4o-generated data and testing on RWD for a binary prediction task showed an F1 score (0.706) with comparable performance to training on the CTGAN data (0.705) for predicting postoperative functional status deterioration. GPT-4o demonstrated a promising ability to generate high-fidelity synthetic neurosurgical data. These findings also indicate that data synthesized with GPT-4o can effectively augment clinical data with small sample sizes, and train ML models for prediction of neurosurgical outcomes. Further investigation is necessary to improve the preservation of distributional characteristics and boost classifier performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。