通过图文定位精准评估肺部影像报告生成质量
Evaluating Automated Radiology Report Quality through Fine-Grained Phrasal Grounding of Clinical Findings
- 提取病灶位置、侧别和严重程度的细粒度模式
- 将文本描述与图像区域匹配,实现图文联合评估
- 能有效检测事实性错误,适合医学AI报告质检
近年来,已有多种评估指标通过词汇、语义或临床命名实体识别方法,仅基于文本信息自动评估胸部X光片生成报告的质量。本文提出一种新方法:首先提取大量临床发现的位置、侧别和严重程度等细粒度发现模式,再通过短语定位(phrasal grounding)将这些发现与胸片图像中的解剖区域对应。结合文本与视觉特征,综合评定生成报告质量。我们在基于MIMIC数据集构建的黄金标准数据集上,将该评估方法与其他纯文本指标进行对比,结果表明其对事实性错误具有更强的敏感性和鲁棒性。
原文摘要 · Abstract (English)
Several evaluation metrics have been developed recently to automatically assess the quality of generative AI reports for chest radiographs based only on textual information using lexical, semantic, or clinical named entity recognition methods. In this paper, we develop a new method of report quality evaluation by first extracting fine-grained finding patterns capturing the location, laterality, and severity of a large number of clinical findings. We then performed phrasal grounding to localize their associated anatomical regions on chest radiograph images. The textual and visual measures are then combined to rate the quality of the generated reports. We present results that compare this evaluation metric with other textual metrics on a gold standard dataset derived from the MIMIC collection and show its robustness and sensitivity to factual errors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。