AI生成医疗数据正侵蚀病理多样性,导致诊断可靠性下降。
AI-generated data contamination erodes pathological variability and diagnostic reliability
- AI生成数据自我强化,导致罕见病征消失、性别年龄偏倚
- 模型误报率翻倍至40%,却仍给出高置信度假性安心报告
- 医生盲评证实:两代后生成文档已无临床价值
生成式人工智能正快速填充医疗记录中的合成内容,形成未来模型训练依赖未经审核的AI生成数据的反馈循环。我们分析超过80万条合成数据,涵盖临床文本生成、视觉语言报告与医学图像合成,发现无论模型架构如何,其输出均趋向通用表型。罕见但关键的异常(如气胸、积液)在生成内容中逐渐消失,人口统计学分布严重偏向中年男性。更严重的是,模型错误地保持高诊断置信度,假性安心率提升至40%。盲法医生评估确认,信心与准确率脱钩,仅两代后生成文档即丧失临床价值。三种缓解策略测试表明,单纯扩大合成数据量无效,而结合真实数据与质量感知过滤可有效维持多样性。研究提示:若无政策强制的人类监督,生成式AI将破坏其赖以生存的医疗数据生态。
原文摘要 · Abstract (English)
Generative artificial intelligence (AI) is rapidly populating medical records with synthetic content, creating a feedback loop where future models are increasingly at risk of training on uncurated AI-generated data. However, the clinical consequences of this AI-generated data contamination remain unexplored. Here, we show that in the absence of mandatory human verification, this self-referential cycle drives a rapid erosion of pathological variability and diagnostic reliability. By analysing more than 800,000 synthetic data points across clinical text generation, vision-language reporting, and medical image synthesis, we find that models progressively converge toward generic phenotypes regardless of the model architecture. Specifically, rare but critical findings, including pneumothorax and effusions, vanish from the synthetic content generated by AI models, while demographic representations skew heavily toward middle-aged male phenotypes. Crucially, this degradation is masked by false diagnostic confidence; models continue to issue reassuring reports while failing to detect life-threatening pathology, with false reassurance rates tripling to 40%. Blinded physician evaluation confirms that this decoupling of confidence and accuracy renders AI-generated documentation clinically useless after just two generations. We systematically evaluate three mitigation strategies, finding that while synthetic volume scaling fails to prevent collapse, mixing real data with quality-aware filtering effectively preserves diversity. Ultimately, our results suggest that without policy-mandated human oversight, the deployment of generative AI threatens to degrade the very healthcare data ecosystems it relies upon.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。