arXiv:2602.04617cs.CL2026-02

通过分层专家对齐,减少医学影像报告生成中的幻觉。

LEAD: Layer-wise Expert-aligned Decoding for Faithful Radiology Report Generation

  • 在每层解码器中引入专家特征,动态修正生成偏差。
  • 在多个公开数据集上显著降低幻觉率,提升临床准确率。
  • 适合需要高可信度医疗报告生成的场景。

放射科报告生成(RRG)旨在从医学影像中生成准确且连贯的诊断内容。尽管大型视觉语言模型(LVLM)提升了报告的流畅性和准确性,但仍存在幻觉问题,即生成看似合理却与图像无关的病理细节。现有方法多依赖外部知识引导以实现文本与视觉信息对齐,但常忽略预训练模型中的固有解码先验和视觉-语言对齐偏差,且因依赖人工构建的引导而缺乏鲁棒性。本文提出分层专家对齐解码(LEAD),一种内生式修改LVLM解码轨迹的新方法。设计多专家模块提取不同病理特征,并通过门控机制融入每一解码层。该分层结构使大语言模型在每个推理步骤中可借助学习到的门控函数调用专家特征,动态纠正解码偏差,引导生成更符合事实的内容。在多个公开数据集上的实验表明,LEAD方法在提升临床准确率的同时有效缓解了幻觉问题,且保持了高水平的生成质量。

原文摘要 · Abstract (English)

Radiology Report Generation (RRG) aims to produce accurate and coherent diagnostics from medical images. Although large vision language models (LVLM) improve report fluency and accuracy, they exhibit hallucinations, generating plausible yet image-ungrounded pathological details. Existing methods primarily rely on external knowledge guidance to facilitate the alignment between generated text and visual information. However, these approaches often ignore the inherent decoding priors and vision-language alignment biases in pretrained models and lack robustness due to reliance on constructed guidance. In this paper, we propose Layer-wise Expert-aligned Decoding (LEAD), a novel method to inherently modify the LVLM decoding trajectory. A multiple experts module is designed for extracting distinct pathological features which are integrated into each decoder layer via a gating mechanism. This layer-wise architecture enables the LLM to consult expert features at every inference step via a learned gating function, thereby dynamically rectifying decoding biases and steering the generation toward factual consistency. Experiments conducted on multiple public datasets demonstrate that the LEAD method yields effective improvements in clinical accuracy metrics and mitigates hallucinations while preserving high generation quality.

医学报告幻觉抑制视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。