arXiv:2511.23066cs.CVcs.AI2025-11

用生成模型修复骨龄影像后,反而让AI诊断更不准。

Evaluating the Clinical Impact of Generative Inpainting on Bone Age Estimation

  • 用自然语言指令让大模型修复手部X光中的非解剖标记
  • 骨龄预测误差从6.26月升至30.11月,性别分类准确率下降
  • 修复后图像出现像素偏差,可能隐藏临床关键特征

生成式基础模型可通过真实感图像修复去除视觉伪影,但其对医学AI性能的影响尚不明确。儿科手部X光常含非解剖标记,目前尚不清楚修复这些区域是否会保留骨龄与性别预测所需特征。为评估生成模型修复在临床中的可靠性,研究使用RSNA骨龄挑战数据集,选取200张原始影像,通过gpt-image-1模型生成600张修复版本,利用自然语言提示定位非解剖伪影。采用深度学习集成模型评估骨龄估计与性别分类表现,以平均绝对误差(MAE)和受试者工作特征曲线下面积(AUC)为指标,并分析像素强度分布以检测结构变化。结果显示,修复后模型性能显著下降:骨龄预测的MAE由6.26月增至30.11月,性别分类的AUC由0.955降至0.704。修复图像呈现像素强度偏移与不一致,表明存在未被简单校准的结构性改变。研究证实,尽管视觉上逼真,基于基础模型的修复可能遮蔽细微但临床重要的特征,并在仅编辑非诊断区域时引入潜在偏差,强调必须进行任务特异性验证,才能将此类生成工具纳入临床AI流程。

原文摘要 · Abstract (English)

Generative foundation models can remove visual artifacts through realistic image inpainting, but their impact on medical AI performance remains uncertain. Pediatric hand radiographs often contain non-anatomical markers, and it is unclear whether inpainting these regions preserves features needed for bone age and gender prediction. To evaluate the clinical reliability of generative model-based inpainting for artifact removal, we used the RSNA Bone Age Challenge dataset, selecting 200 original radiographs and generating 600 inpainted versions with gpt-image-1 using natural language prompts to target non-anatomical artifacts. Downstream performance was assessed with deep learning ensembles for bone age estimation and gender classification, using mean absolute error (MAE) and area under the ROC curve (AUC) as metrics, and pixel intensity distributions to detect structural alterations. Inpainting markedly degraded model performance: bone age MAE increased from 6.26 to 30.11 months, and gender classification AUC decreased from 0.955 to 0.704. Inpainted images displayed pixel-intensity shifts and inconsistencies, indicating structural modifications not corrected by simple calibration. These findings show that, although visually realistic, foundation model-based inpainting can obscure subtle but clinically relevant features and introduce latent bias even when edits are confined to non-diagnostic regions, underscoring the need for rigorous, task-specific validation before integrating such generative tools into clinical AI workflows.

医学AI生成修复骨龄评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。