arXiv:2602.23676cs.CV2026-02被引 1

用语义解耦方法抑制影像报告生成中的历史幻觉,提升临床准确性。

Suppressing Prior-Comparison Hallucinations in Radiology Report Generation via Semantically Decoupled Latent Steering

  • 通过大语言模型分解与正交化构建无语义干扰的控制向量
  • 在MIMIC-CXR上将历史幻觉概率降低20.4%,临床标签准确率提升43%
  • 无需训练,可直接用于现有模型,适合医疗AI开发者

基于视觉-语言模型的自动化放射科报告生成受限于先验对比幻觉问题,即模型生成当前检查中未支持的历史发现。本文提出无需训练、仅在推理时生效的语义解耦潜在空间引导(SDLS)框架。不同于通用激活引导常存在语义纠缠,本方法利用大语言模型进行语义分解,并通过QR正交化构建语义无关的干预向量。该正交化步骤至关重要,其借助几何约束过滤标准主成分分析方向中常见的临床语义,确保引导向量仅作用于“历史对比”维度。在BiomedGPT基础模型上验证表明,该方法克服了幻觉抑制与临床准确性之间的权衡。在MIMIC-CXR上的大量实验及在CheXpert Plus和IU-Xray上的零样本迁移评估均证明其鲁棒性。定量分析显示,该方法显著降低历史幻觉概率(FilBERT得分从0.2373降至0.1889),提升临床标签保真度(CheXpert宏F1从0.2242升至0.3208)。附加评估确认临床叙述结构完整性得以保持。

原文摘要 · Abstract (English)

Automated radiology report generation using vision-language models (VLMs) is limited by the risk of prior-comparison hallucination, where the model generates historical findings unsupported by the current study. We address this challenge with a training-free, inference-time control framework termed Semantically Decoupled Latent Steering (SDLS). Unlike generic activation steering, which often suffers from semantic entanglement, our approach constructs a semantic-free intervention vector via large language model (LLM)-driven semantic decomposition followed by $QR$-based orthogonalization. This orthogonalization step is critical. It leverages geometric constraints to filter out the clinical semantics often entangled in standard principal component analysis (PCA) directions, ensuring that the steering vector targets only the ``historical comparison" axis. We validate our method on the BiomedGPT foundation model, demonstrating that it overcomes the trade-off between hallucination suppression and clinical accuracy. Extensive experiments on MIMIC-CXR, and zero-shot transfer evaluation on CheXpert Plus and IU-Xray, demonstrate the robustness of our approach. Quantitative evaluations on MIMIC-CXR show that our approach significantly reduces the probability of historical hallucinations (FilBERT score decreases from 0.2373 to 0.1889) and improves clinical label fidelity (CheXpert macro-F1 increases from 0.2242 to 0.3208). Supplementary evaluations confirm that the structural integrity of the clinical narrative is maintained.

医学报告生成幻觉抑制视觉语言模型生成控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。