arXiv:2509.21356cs.CVcs.AI2025-09被引 4

用合成数据训练模型,精准检测肺部X光报告中的事实错误和位置错配。

Phrase-grounded Fact-checking for Automatically Generated Chest X-ray Reports

  • 基于真实报告生成带错的合成数据,模拟错误发现与位置匹配问题。
  • 在多个数据集上实现0.997的校验一致性,接近人工标注精度。
  • 适合用于顶尖报告生成模型的错误检测,提升临床报告可信度。

随着大规模视觉语言模型(VLM)的发展,现在可以为胸部X光图像生成逼真的放射科报告。然而,这些报告在推理过程中存在事实性错误和幻觉,阻碍了其临床应用。本文提出一种新的短语锚定式事实核查模型(FC模型),用于检测自动生成的胸部放射科报告中发现内容及其对应位置的错误。具体而言,我们通过扰动真实报告中的发现及其位置,构建一个大规模合成数据集,形成真实与虚假的发现-位置配对。随后,在该数据集上训练一种新型多标签跨模态对比回归网络。实验结果表明,该方法在多个X光数据集上均表现出优异的真伪判断准确率与定位性能。同时,其在多个数据集上对当前最先进报告生成器输出的报告进行错误检测时,与基于真实标注的验证相比,达到了0.997的组内相关系数,显示出其在放射科临床推理流程中的实用价值。

原文摘要 · Abstract (English)

With the emergence of large-scale vision language models (VLM), it is now possible to produce realistic-looking radiology reports for chest X-ray images. However, their clinical translation has been hampered by the factual errors and hallucinations in the produced descriptions during inference. In this paper, we present a novel phrase-grounded fact-checking model (FC model) that detects errors in findings and their indicated locations in automatically generated chest radiology reports. Specifically, we simulate the errors in reports through a large synthetic dataset derived by perturbing findings and their locations in ground truth reports to form real and fake findings-location pairs with images. A new multi-label cross-modal contrastive regression network is then trained on this dataset. We present results demonstrating the robustness of our method in terms of accuracy of finding veracity prediction and localization on multiple X-ray datasets. We also show its effectiveness for error detection in reports of SOTA report generators on multiple datasets achieving a concordance correlation coefficient of 0.997 with ground truth-based verification, thus pointing to its utility during clinical inference in radiology workflows.

医学报告事实核查视觉语言模型肺部影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。