首个统一评估合成胸片质量、隐私与实用性的基准,填补医疗生成领域空白。
CheXGenBench: A Unified Benchmark For Fidelity, Privacy and Utility of Synthetic Chest Radiographs
- 构建覆盖11种主流模型的统一评估框架,支持新模型快速接入。
- 顶尖模型仍难处理罕见病分布,且高保真不等于低隐私风险。
- 合成数据对分类任务有用,但对多模态任务帮助有限,适合医疗生成研究者参考。
结构化基准已推动真实图像的文本条件生成发展,但针对合成胸片生成尚无此类标准。尽管该领域研究活跃,现有工作仍采用不一致的评估方式,缺乏对生成保真度、隐私风险和下游实用性三大核心指标的统一评估。为此,我们提出CheXGenBench,首个统一评估合成胸片生成的基准框架,可同时评估前沿文本到图像(T2I)生成模型在保真度、隐私风险与下游实用性方面的表现。评估协议包含20余项定量指标,覆盖11个主流T2I架构,支持新模型即插即用。通过严谨公平的评测,我们建立了各维度的综合最先进(SoTA)基线性能,为未来研究提供指引。此外,结果揭示当前模型存在三大局限:第一,即使最先进模型也难以应对长尾医学分布;第二,无论保真度高低,模型均存在显著隐私风险;第三,尽管合成数据已能提升下游分类性能,但在多模态任务中效用有限。基于此,我们提出了具体的研究方向以推动该领域发展。代码已开源:https://github.com/Raman1121/CheXGenBench。
原文摘要 · Abstract (English)
Structured benchmarks have advanced text-conditional image generation for real-world imagery, however, no such benchmark exists for synthetic radiograph generation. Despite being a highly active area of research, existing studies continue adopting inconsistent evaluation protocols and lack a unified assessment of the three most critical criteria: generative fidelity, privacy risk, and downstream utility. To address these limitations, we introduce CheXGenBench, the first unified evaluation framework for synthetic chest radiograph generation that simultaneously assesses fidelity, privacy risks, and downstream utility across frontier text-to-image (T2I) generative models. Our evaluation protocol, comprising over 20 quantitative metrics, covers 11 leading T2I architectures with plug-and-play integration for newer models. Through a rigorous and fair evaluation protocol, we establish comprehensive baseline state-of-the-art (SoTA) performances across all dimensions to guide future research. Furthermore, our results uncover several limitations of current generative models, which include first, even SoTA models struggle with long-tailed medical distributions; second, models pose high privacy risks regardless of fidelity quality; and third, while synthetic data already benefits downstream classification, it is of limited utility for downstream multimodal tasks. Drawing from these results, we propose concrete research directions to advance the field. The code is available at https://github.com/Raman1121/CheXGenBench
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。