提出诊断链框架,让放射报告生成更准确且可解释。
A Chain of Diagnosis Framework for Accurate and Explainable Radiology Report Generation
- 用问答对话提取关键病灶,再由大模型生成报告。
- 在两个基准上优于专业与通用模型,病灶定位更准。
- 适合需要可信报告的临床医生和研究者使用。
尽管放射科报告生成(RRG)取得进展,现有方法仍面临两大挑战:1)临床有效性不足,尤其在病灶特征描述上;2)生成文本缺乏可解释性,难获放射科医生信任。为此,本文提出可信赖的RRG框架——诊断链(CoD),实现临床准确且可解释的报告生成。该框架首先通过诊断对话生成问题-答案对以提取关键发现,再利用大语言模型根据问答结果生成报告。为增强可解释性,设计诊断定位模块,将生成语句与问答诊断进行匹配;同时引入病灶定位模块,在图像中精确定位异常区域,提升放射科医生的工作效率。为支持高效训练,提出融合临床一致性的全监督学习策略,利用多源标注数据。本研究构建了包含问答对与病灶框的全标注RRG数据集,开发了评估报告中病灶位置与严重程度描述准确性的工具,并在两个基准上验证了CoD的有效性:其性能持续优于专精与通用模型,且能精准将生成句子映射至问答诊断和图像区域,展现良好可解释性。
原文摘要 · Abstract (English)
Despite the progress of radiology report generation (RRG), existing works face two challenges: 1) The performances in clinical efficacy are unsatisfactory, especially for lesion attributes description; 2) the generated text lacks explainability, making it difficult for radiologists to trust the results. To address the challenges, we focus on a trustworthy RRG model, which not only generates accurate descriptions of abnormalities, but also provides basis of its predictions. To this end, we propose a framework named chain of diagnosis (CoD), which maintains a chain of diagnostic process for clinically accurate and explainable RRG. It first generates question-answer (QA) pairs via diagnostic conversation to extract key findings, then prompts a large language model with QA diagnoses for accurate generation. To enhance explainability, a diagnosis grounding module is designed to match QA diagnoses and generated sentences, where the diagnoses act as a reference. Moreover, a lesion grounding module is designed to locate abnormalities in the image, further improving the working efficiency of radiologists. To facilitate label-efficient training, we propose an omni-supervised learning strategy with clinical consistency to leverage various types of annotations from different datasets. Our efforts lead to 1) an omni-labeled RRG dataset with QA pairs and lesion boxes; 2) a evaluation tool for assessing the accuracy of reports in describing lesion location and severity; 3) extensive experiments to demonstrate the effectiveness of CoD, where it outperforms both specialist and generalist models consistently on two RRG benchmarks and shows promising explainability by accurately grounding generated sentences to QA diagnoses and images.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。