arXiv:2607.25589cs.CVcs.AI2026-07

复现审计发现医学AI基准存在多处数据偏差,影响结论可信度。

Forensic Reproducibility Audit of a Radiology Vision-Language Model Benchmark: From Intended Protocol to Released Artifact

  • 回溯审计原模型调用与输出,验证各环节一致性
  • 300次调用中297次生成非空报告,但部分图像渲染出错
  • 质疑原始性能结论,提出可机器验证的控制标准

医疗影像AI基准包含数据集、DICOM渲染、提示词、接口、自动标签、统计代码、论文和代码库发布。这些组件间的统一性通常被默认而非验证。我们对一个保留下来的胸部X光视觉语言模型(VLM)试点项目进行了回溯性复现审计,未重新调用模型,也未新标注图像或报告。追踪了提示词绑定、DICOM元数据、输出完整性、标签提取、匹配分析及发布传播过程。300次计划中的模型-提示调用中,297次产生非空报告。60次使用相同提示词的Claude调用生成了A/B两组结果。30项研究对应28名患者。4张MONOCHROME1图像未执行必需的极性反转,数据集划分归属未保留,未经验证的提取器将5份报告截断至4000字符。重建369个完整病例发现块后,Cochran's Q值从154.73升至182.29。在45次McNemar检验中,27次未校正p<0.05,20次经Holm校正后仍低于0.05。这些数值仅反映存档自动标签矩阵;无法恢复原定提示对比,也未建立临床性能。我们撤回原有性能、排名、提示效应和临床声明,并明确机器可验证的控制要求:队列、DICOM渲染、提示与模型身份、调用状态、标注溯源、关键分析及衍生产物。

原文摘要 · Abstract (English)

Medical-imaging AI benchmarks combine datasets, DICOM rendering, prompts, provider APIs, automated labels, statistical code, manuscripts, and repository releases. Agreement across these artifacts is usually assumed rather than tested. We performed a retrospective forensic reproducibility audit of a preserved chest-radiograph vision-language model (VLM) pilot; no model was called again and no image or report was newly annotated. We traced prompt bindings, DICOM metadata, output completeness, label extraction, matched analyses, and release propagation. Of 300 planned model-prompt calls, 297 yielded nonempty reports. Sixty Claude calls labeled A/B were executed with the same C prompt. The 30 studies represented 28 patients. Four MONOCHROME1 images were rendered without required polarity inversion, dataset split membership was not retained, and the unvalidated extractor truncated five reports to 4000 characters. Reconstructing one common cohort of 369 complete case-finding blocks changed Cochran's Q from 154.73 to 182.29. Of 45 McNemar comparisons, 27 had unadjusted p < 0.05 and 20 remained below 0.05 after Holm adjustment. These values describe only the archived automated-label matrix; they do not recover the intended prompt comparison or establish clinical performance. We withdraw the original performance, ranking, prompt-effect, and clinical claims and specify machine-verifiable controls for cohort, DICOM rendering, prompt and model identity, call status, annotation provenance, keyed analysis, and derived artifacts.

医学AI复现审计视觉语言模型数据偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。