无需真实病历生成高保真合成患者,突破隐私与数据偏差瓶颈
Patient-Zero: Scaling Synthetic Patient Agents to Real-World Distributions without Real Patient Data
- 从临床指南出发,分层排列属性生成多样化患者记录
- 人类医生评估认为合成数据与真实数据无统计差异,临床质量更高
- 下游医学推理模型性能提升显著,适合医疗AI训练与测试
大型语言模型生成合成数据在医疗领域成为缓解数据稀缺与隐私限制的有前景方案。然而,现有方法仍依赖真实病历,存在隐私风险和分布偏差。此外,当前患者代理面临稳定性与可塑性矛盾,难以在动态交互中保持临床一致性。为此,我们提出Patient-Zero框架,无需真实医疗记录即可从头构建患者模拟。其医学对齐的分层合成机制通过分层属性排列,从抽象临床指南生成全面且多样的患者记录。为支持严谨临床交互,设计双轨认知记忆系统,实现动态记忆更新的同时维持逻辑一致性和角色一致性。大量评估显示,Patient-Zero在数据质量和交互保真度上达到新基准。在人类专家评估中,资深持证医师认为其合成数据在统计上无法与真人撰写数据区分,且临床质量更优。此外,基于该合成数据集训练的下游医学推理模型性能显著提升:MedQA +24.0%,MMLU +14.5%,验证了框架的实际价值。
原文摘要 · Abstract (English)
Synthetic data generation with Large Language Models (LLMs) has emerged as a promising solution in the medical domain to mitigate data scarcity and privacy constraints. However, existing approaches remain constrained by their derivative nature, relying on real-world records, which pose privacy risks and distribution biases. Furthermore, current patient agents face the Stability-Plasticity Dilemma, struggling to maintain clinical consistency during dynamic inquiries. To address these challenges, we introduce Patient-Zero, a novel framework for ab initio patient simulation that requires no real medical records. Our Medically-Aligned Hierarchical Synthesis framework generates comprehensive and diverse patient records from abstract clinical guidelines via stratified attribute permutation. To support rigorous clinical interaction, we design a Dual-Track Cognitive Memory System to enable agents dynamically update memory while preserving logical consistency and persona adherence. Extensive evaluations show that Patient-Zero establishes a new state-of-the-art in both data quality and interaction fidelity. In human expert evaluations, senior licensed physicians judge our synthetic data to be statistically indistinguishable from real human-authored data and higher in clinical quality. Furthermore, downstream medical reasoning model trained on our synthetic dataset shows substantial performance gains (MedQA +24.0%; MMLU +14.5%), demonstrating the practical utility of our framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。