arXiv:2507.21123cs.AIcs.LG2025-07

用大模型加速生成更真实、多样的病患数据模块

Leveraging Generative AI to Enhance Synthea Module Development

  • 用大模型自动生成疾病描述和数据模块
  • 通过迭代评估与修正提升模块准确性和语法正确性
  • 适合医疗数据生成研究者快速构建新疾病模型

本文探讨了利用大语言模型(LLMs)辅助Synthea开源健康数据生成器开发新疾病模块的可行性。将LLMs融入模块开发流程,有望缩短开发时间、降低专业门槛、丰富模型多样性并提升合成患者数据的整体质量。文中展示了四种应用场景:生成疾病特征描述、基于描述生成模块、评估现有模块以及优化已有模块。提出渐进式优化概念,即通过持续检查生成模块的语法正确性与临床合理性,并据此迭代修改。尽管该方法前景广阔,但仍需人工监督、严格测试验证,并警惕生成内容的潜在偏差。论文最后提出未来研究建议,以充分发挥大模型在合成数据生成中的潜力。

原文摘要 · Abstract (English)

This paper explores the use of large language models (LLMs) to assist in the development of new disease modules for Synthea, an open-source synthetic health data generator. Incorporating LLMs into the module development process has the potential to reduce development time, reduce required expertise, expand model diversity, and improve the overall quality of synthetic patient data. We demonstrate four ways that LLMs can support Synthea module creation: generating a disease profile, generating a disease module from a disease profile, evaluating an existing Synthea module, and refining an existing module. We introduce the concept of progressive refinement, which involves iteratively evaluating the LLM-generated module by checking its syntactic correctness and clinical accuracy, and then using that information to modify the module. While the use of LLMs in this context shows promise, we also acknowledge the challenges and limitations, such as the need for human oversight, the importance of rigorous testing and validation, and the potential for inaccuracies in LLM-generated content. The paper concludes with recommendations for future research and development to fully realize the potential of LLM-aided synthetic data creation.

生成模型医疗数据AI辅助

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。