用AI生成医疗数据,隐私安全又高质量。
Leveraging Generative AI Through Prompt Engineering and Rigorous Validation to Create Comprehensive Synthetic Datasets for AI Training in Healthcare
- 用GPT-4和提示工程生成完整病历数据。
- 通过5种模型验证,确保数据真实可信。
- 适合需训练医疗AI但难获真实数据的团队。
由于隐私问题,获取高质量医疗数据常受限制,严重制约电子病历(EHR)应用中人工智能算法的训练。本研究采用GPT-4 API进行提示工程,生成涵盖患者入院信息全维度的合成数据,包括医疗机构信息、科室、病房、床位、人口统计、紧急联系人、生命体征、疫苗接种、过敏史、病史、预约、就诊记录、检验项目、诊断、治疗方案、用药、临床笔记、访视日志、出院小结及转诊信息。为保障数据质量,引入多种验证技术:使用BERT的下一句预测评估句子连贯性,GPT-2判断整体合理性,RoBERTa检测逻辑一致性,自编码器识别异常值,并开展多样性分析。通过所有验证标准的数据被整合至基于PostgreSQL的综合数据库,作为EHR应用的数据管理系统。结果表明,结合生成式AI与严格验证可有效生成高质量医疗合成数据,助力AI训练同时规避真实患者数据的隐私风险。
原文摘要 · Abstract (English)
Access to high-quality medical data is often restricted due to privacy concerns, posing significant challenges for training artificial intelligence (AI) algorithms within Electronic Health Record (EHR) applications. In this study, prompt engineering with the GPT-4 API was employed to generate high-quality synthetic datasets aimed at overcoming this limitation. The generated data encompassed a comprehensive array of patient admission information, including healthcare provider details, hospital departments, wards, bed assignments, patient demographics, emergency contacts, vital signs, immunizations, allergies, medical histories, appointments, hospital visits, laboratory tests, diagnoses, treatment plans, medications, clinical notes, visit logs, discharge summaries, and referrals. To ensure data quality and integrity, advanced validation techniques were implemented utilizing models such as BERT's Next Sentence Prediction for sentence coherence, GPT-2 for overall plausibility, RoBERTa for logical consistency, autoencoders for anomaly detection, and conducted diversity analysis. Synthetic data that met all validation criteria were integrated into a comprehensive PostgreSQL database, serving as the data management system for the EHR application. This approach demonstrates that leveraging generative AI models with rigorous validation can effectively produce high-quality synthetic medical data, facilitating the training of AI algorithms while addressing privacy concerns associated with real patient data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。