arXiv:2608.07550cs.CVcs.LG2026-08

评估医学影像-文本模型在不同机构间的判断一致性,发现需按机构和接口单独校准信任度。

Auditing Medical Vision-Language Models on Chest Radiographs: Estimating Reference Agreement Across Institutions

论文配图:Auditing Medical Vision-Language Models on Chest Radiographs: Estimating Reference Agreement Across Institutions
图 1 · 摘自论文原文
  • 通过小样本本地标注估算各机构的判断一致性,避免跨机构直接迁移。
  • 三种模型在34.5万条预测中表现接近,最优估计器仅比基准略优0.0003。
  • 结果表明信任度必须按机构和提问方式重新评估,不能通用。

视觉语言模型通过不显示置信度的接口返回胸部X光片的结构化结论,接收机构无法直接判断其可信程度。跨机构、不同病灶、预测方向及问题形式下与机构参考标准的一致性尚未被量化测量。我们对三个生成式视觉语言模型在三个机构的胸片数据集上进行了评估,涵盖六种病灶类型和两种提问方式,共超过34.5万次病灶级预测,并基于少量本地标签估算接收机构的病灶-方向级参考一致性。估计策略在多次严格排除接收机构的机构保留评估中被压力测试。当完全排除接收机构参与开发时,七种估计器中自适应选择并未优于固定方法:均方误差(Brier分数)为0.1083,而始终使用贝塔-二项分布经验贝叶斯估计器得0.0853,仅目标项逻辑回归模型得0.0855。两者差距仅为0.0003,小于该方法对求解器版本变化的敏感度,且各在约一半场景中领先,故无法推荐默认方案。二者相对于跨机构聚合估计器的优势集中于单一机构,且在按机构聚类后不具可确认性;名义95%的插值经验贝叶斯后验预测区间覆盖率为87.0%,在最难机构更低。因此,参考一致性必须按机构和接口重新评估,此研究关注的是与机构标签的一致性,而非临床正确性。

原文摘要 · Abstract (English)

Vision-language models return structured chest-radiograph findings through interfaces exposing no confidence score, so a receiving institution cannot read off how far to trust an individual judgment. Whether agreement with an institution's reference standard transfers across sites, findings, prediction directions and question formats is largely unmeasured. We evaluated three generative vision-language models on three institutional chest-radiograph corpora and six findings under two elicitation protocols, comprising more than 345,000 finding-level predictions, and estimated finding-by-direction reference agreement at a receiving institution from a small budget of local labels. Estimation strategies were then stress-tested under repeated strict institution-held-out evaluation. Under evaluation excluding the receiving institution from development entirely, adaptive selection among the seven estimators that design admits did not improve on simple fixed alternatives: it achieved a mean Brier score of 0.1083, against 0.0853 for always using a Beta-Binomial empirical-Bayes estimator and 0.0855 for a target-only logistic model. Those two differ by 0.0003, less than this family's own sensitivity to a change of solver version, and each leads in about half the settings, so no default can be recommended. Their advantage over estimators pooling across institutions was concentrated at one site and not confirmatory once clustered by institution, and a plug-in empirical-Bayes posterior-predictive count interval at a nominal 95% level covered 87.0%, less at the hardest institution. Reference agreement therefore has to be re-evaluated per site and per interface; these results concern agreement with institutional labels, not clinical correctness.

医学影像模型审计一致性评估多中心

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。