用大模型自动生成医学影像报告,确保内容准确不捏造。
LLM-Bootstrapped Targeted Finding Guidance for Factual MLLM-based Medical Report Generation
- 先用大模型识别影像中的真实病灶,再引导生成报告
- 在两个疾病数据集上准确率显著优于现有方法
- 无需人工标注,可自动构建带标签的训练数据
利用多模态大语言模型(MLLM)自动生成医学报告时,常因事实不稳导致遗漏发现或引入错误信息,限制了其临床应用。现有方法直接基于图像特征生成报告,缺乏明确的事实依据。为此,我们提出Fact-Flow框架,将视觉事实识别与报告生成分离:先通过图像预测临床发现,再引导MLLM生成精准报告。关键创新在于采用大语言模型(LLM)自主构建带标签的医学发现数据集,避免昂贵的人工标注。在两个疾病聚焦的医疗数据集上的广泛实验表明,该方法在显著提升事实准确性的同时,保持了高水平的文本质量,优于当前最优模型。
原文摘要 · Abstract (English)
The automatic generation of medical reports utilizing Multimodal Large Language Models (MLLMs) frequently encounters challenges related to factual instability, which may manifest as the omission of findings or the incorporation of inaccurate information, thereby constraining their applicability in clinical settings. Current methodologies typically produce reports based directly on image features, which inherently lack a definitive factual basis. In response to this limitation, we introduce Fact-Flow, an innovative framework that separates the process of visual fact identification from the generation of reports. This is achieved by initially predicting clinical findings from the image, which subsequently directs the MLLM to produce a report that is factually precise. A pivotal advancement of our approach is a pipeline that leverages a Large Language Model (LLM) to autonomously create a dataset of labeled medical findings, effectively eliminating the need for expensive manual annotation. Extensive experimental evaluations conducted on two disease-focused medical datasets validate the efficacy of our method, demonstrating a significant enhancement in factual accuracy compared to state-of-the-art models, while concurrently preserving high standards of text quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。