arXiv:2411.19378cs.CVcs.AI2024-11ACL被引 35

Libra通过时间对齐机制提升胸片报告生成的临床准确性

Libra: Leveraging Temporal Images for Biomedical Radiology Analysis

  • 用专用影像编码器+时间对齐连接器捕捉前后影像差异
  • 在MIMIC-CXR数据集上超越同类模型,临床相关性与词汇准确率双高
  • 适合需要时序分析能力的医学影像报告生成任务

放射科报告生成(RRG)需具备先进的医学图像分析、有效的时序推理和精准的文本生成能力。尽管多模态大模型(MLLMs)通过预训练视觉编码器增强了视觉-语言理解,但现有方法大多依赖单图分析或基于规则的启发式方法处理多张图像,未能充分利用多模态医疗数据中的时序信息。本文提出Libra,一种面向胸部X光报告生成的时间感知型MLLM。Libra结合专用于放射科的图像编码器与创新的时间对齐连接器(TAC),能够准确捕捉并整合配对当前与既往影像间的时序差异。在MIMIC-CXR数据集上的大量实验表明,Libra在同等规模的MLLM中建立了新的最先进基准,显著提升了临床相关性与词汇准确率。

原文摘要 · Abstract (English)

Radiology report generation (RRG) requires advanced medical image analysis, effective temporal reasoning, and accurate text generation. While multimodal large language models (MLLMs) align with pre-trained vision encoders to enhance visual-language understanding, most existing methods rely on single-image analysis or rule-based heuristics to process multiple images, failing to fully leverage temporal information in multi-modal medical datasets. In this paper, we introduce Libra, a temporal-aware MLLM tailored for chest X-ray report generation. Libra combines a radiology-specific image encoder with a novel Temporal Alignment Connector (TAC), designed to accurately capture and integrate temporal differences between paired current and prior images. Extensive experiments on the MIMIC-CXR dataset demonstrate that Libra establishes a new state-of-the-art benchmark among similarly scaled MLLMs, setting new standards in both clinical relevance and lexical accuracy.

医学影像时序分析报告生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。