arXiv:2605.05761cs.CV2026-05被引 1

构建可编程虚拟病灶实验框架,精准评估肺CT模型性能

iTRIALSPACE: Programmable Virtual Lesion Trials for Controlled Evaluation of Lung CT Models

论文配图:iTRIALSPACE: Programmable Virtual Lesion Trials for Controlled Evaluation of Lung CT Models
图 1 · 摘自论文原文
  • 通过四阶段流程合成可控虚拟病灶数据,分离真实影像与病灶特征
  • 13种实验模式下合成数据与真实数据相似度高,性能排名与真实数据高度一致
  • 适合需要可重复、可验证模型评估的研究者,尤其关注病灶大小误判等偏差

我们提出iTRIALSPACE,一个可编程的评估框架,用于对肺CT模型进行受控测试。传统基准是静态回顾性数据集,混杂了病灶大小、肺叶分布、解剖结构和成像条件,难以判断影响模型准确性的具体因素。iTRIALSPACE通过四阶段流程:多数据源结节特征分析、显式试验设定、解剖感知掩码插入和ControlNet条件化图像生成,将真实临床CT与结节特征组合成可控的虚拟病灶试验。该框架基于涵盖7个公开CT来源共13,140个标注结节的统一54属性结节特征数据集,实现13种试验模式。我们在包含55,469样本的虚拟病灶研究中评估了该框架,覆盖三种医学视觉语言模型、四种空间引导条件和三种临床任务。所有13种模式下,合成数据的FID值保持在真实数据之间水平,合成性能排名与真实数据强相关(ρ = 0.93,p < 10⁻¹⁵)。受控试验揭示了固定分布基准无法发现的现象,包括在肺叶均衡采样下病灶大小预测出现退化现象,以及双胞胎交叉分析中宿主-供体方差比分别为8.9倍和3.3倍。这些结果表明,iTRIALSPACE可作为超越静态回顾性基准的可审计、可证伪的评估基础设施。

原文摘要 · Abstract (English)

We introduce iTRIALSPACE, a programmable evaluation framework for controlled assessment of lung CT models. Standard benchmarks are static retrospective collections that entangle lesion size, lobe prevalence, anatomy, and acquisition context, making it difficult to determine what structurally drives model accuracy. iTRIALSPACE addresses this limitation by composing real clinical CTs and lesion profiles into controlled virtual lesion trials through a four-stage pipeline: multidataset nodule profiling, explicit trial specification, anatomy-aware mask insertion, and ControlNet-conditioned CT synthesis. The framework is built on a unified 54-attribute nodule-profile dataset spanning 13,140 annotated nodules from seven public CT sources and instantiated as 13 trial modes. We evaluate iTRIALSPACE in a 55,469-sample Virtual Lesion Study spanning three medical VLMs, four spatialguidance conditions, and three clinical tasks. Across all 13 modes, the synthetic substrate remains within the real-to-real FID baseline, and synthetic performance rankings transfer strongly to real clinical data ($ρ$ = 0.93, p < 10$^{-15}$). Controlled trial modes expose findings unavailable to fixed-distribution benchmarks, including shortcut-driven size prediction collapse under lobe-equalized sampling and hostto-donor variance ratios of 8.9x and 3.3x in twin-cross analysis. These results position iTRIALSPACE as an auditable evaluation infrastructure for controlled, falsifiable testing beyond static retrospective benchmarks.

肺CT模型评估虚拟实验可控合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。