arXiv:2410.17357cs.CV2024-10NeurIPS被引 3

新评估指标让机器报告更贴近真实医生判断。

Image-aware Evaluation of Generated Medical Reports

  • 结合图像与报告,综合评估生成报告的临床准确性。
  • 在放射科医生标注错误的数据集上,评分与人工判断高度一致。
  • 提供带精心设计扰动的新数据集,适合评测评估方法优劣。

本文提出一种新型医学影像报告生成评估指标VLScore,旨在克服现有方法的局限性:要么仅关注文本相似度而忽略临床意义,要么只关注病灶识别而忽视其他因素。该指标通过考虑对应图像来衡量报告间的相似性。我们在放射科医生标注错误的成对报告数据集上验证了该方法的有效性,结果显示其评分与医生判断有显著一致性。此外,我们构建了一个新数据集,包含精心设计的扰动,可区分重大修改(如删除诊断)与微小变化,揭示当前评估指标的不足,并为评估方法分析提供清晰框架。

原文摘要 · Abstract (English)

The paper proposes a novel evaluation metric for automatic medical report generation from X-ray images, VLScore. It aims to overcome the limitations of existing evaluation methods, which either focus solely on textual similarities, ignoring clinical aspects, or concentrate only on a single clinical aspect, the pathology, neglecting all other factors. The key idea of our metric is to measure the similarity between radiology reports while considering the corresponding image. We demonstrate the benefit of our metric through evaluation on a dataset where radiologists marked errors in pairs of reports, showing notable alignment with radiologists' judgments. In addition, we provide a new dataset for evaluating metrics. This dataset includes well-designed perturbations that distinguish between significant modifications (e.g., removal of a diagnosis) and insignificant ones. It highlights the weaknesses in current evaluation metrics and provides a clear framework for analysis.

医疗报告图像评估评测指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。