arXiv:2502.09688cs.CVcs.AI2025-02被引 5

用生成模型虚拟临床试验,测试放射AI在真实场景下的表现

Towards Virtual Clinical Trials of Radiology AI with Conditional Generative Modeling

  • 通过条件生成模型合成带属性的全身CT图像
  • 发现AI在真实场景中性能下降最高达20%
  • 适合评估AI模型鲁棒性与偏见,助力临床落地

人工智能有望通过数据驱动洞察重塑医疗,尤其在影像领域。然而,现实应用中AI模型常因泛化能力差导致性能下降高达20%,从受控测试环境到临床实际使用时表现显著下滑。这引发医生被错误预测误导或失去对AI信任的风险,使技术难以落地。为提前识别此类问题,需在多样数据上进行充分临床试验,但真实数据收集与标注成本高昂。为此,我们提出一种新型条件生成模型,用于放射AI的虚拟临床试验(VCT),可真实合成具有指定属性的全身CT图像。通过学习图像与解剖结构的联合分布,该模型能以前所未有的细节重现真实患者群体。基于此合成数据集,我们实现了对放射AI模型的有效评估,揭示了其性能退化,并支持针对诱发偏见的数据属性进行算法审计。该生成式VCT方法为规模化评估模型鲁棒性、缓解偏见、保障患者安全提供了可行路径,使AI模型可在任意目标人群范围内简便测试。

原文摘要 · Abstract (English)

Artificial intelligence (AI) is poised to transform healthcare by enabling personalized and efficient care through data-driven insights. Although radiology is at the forefront of AI adoption, in practice, the potential of AI models is often overshadowed by severe failures to generalize: AI models can have performance degradation of up to 20% when transitioning from controlled test environments to clinical use by radiologists. This mismatch raises concerns that radiologists will be misled by incorrect AI predictions in practice and/or grow to distrust AI, rendering these promising technologies practically ineffectual. Exhaustive clinical trials of AI models on abundant and diverse data is thus critical to anticipate AI model degradation when encountering varied data samples. Achieving these goals, however, is challenging due to the high costs of collecting diverse data samples and corresponding annotations. To overcome these limitations, we introduce a novel conditional generative AI model designed for virtual clinical trials (VCTs) of radiology AI, capable of realistically synthesizing full-body CT images of patients with specified attributes. By learning the joint distribution of images and anatomical structures, our model enables precise replication of real-world patient populations with unprecedented detail at this scale. We demonstrate meaningful evaluation of radiology AI models through VCTs powered by our synthetic CT study populations, revealing model degradation and facilitating algorithmic auditing for bias-inducing data attributes. Our generative AI approach to VCTs is a promising avenue towards a scalable solution to assess model robustness, mitigate biases, and safeguard patient care by enabling simpler testing and evaluation of AI models in any desired range of diverse patient populations.

虚拟临床试验生成模型放射AI模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。