提出可微分多尺度多模态分词器,提升医学影像报告生成质量。
$μ^2$Tokenizer: Differentiable Multi-Scale Multi-Modal Tokenizer for Radiology Report Generation
- 设计多尺度视觉与文本分词融合的可微分分词器
- 在4个大型CT数据集上超越现有方法,报告质量显著提升
- 适合医学AI、临床辅助诊断研究者参考
自动化放射科报告生成(RRG)旨在从临床影像(如CT扫描)中生成详细文本报告,以提高诊断准确性和管理建议效率。该任务面临两大挑战:(1)在资源受限条件下从影像数据中提取相关信息的固有复杂性;(2)模型生成报告与专家撰写报告之间差异的客观评估困难。为此,我们提出μ²LLM——一种用于RRG任务的多尺度多模态大语言模型。其核心组件μ²Tokenizer作为中间层,整合来自多尺度视觉分词器和文本分词器的多模态特征,并通过直接偏好优化(DPO)增强报告生成质量,指导信号来自GREEN-RedLlama。在四个大型CT图像-报告医学数据集上的实验结果表明,该方法优于现有方法,凸显了在有限数据下微调μ²LLM在RRG任务中的潜力。同时,针对提示工程,我们引入一个五阶段、基于大语言模型的流水线,将常规CT报告转化为配对的视觉-问题-答案三元组及带引用的推理叙事,构建可扩展、高质量的监督语料库,用于可解释的多模态放射科大模型训练。所有代码、数据集和模型将在官方仓库公开。
原文摘要 · Abstract (English)
Automated radiology report generation (RRG) aims to produce detailed textual reports from clinical imaging, such as computed tomography (CT) scans, to improve the accuracy and efficiency of diagnosis and provision of management advice. RRG is complicated by two key challenges: (1) inherent complexity in extracting relevant information from imaging data under resource constraints, and (2) difficulty in objectively evaluating discrepancies between model-generated and expert-written reports. To address these challenges, we propose $μ^2$LLM, a $\underline{\textbf{mu}}$ltiscale $\underline{\textbf{mu}}$ltimodal large language models for RRG tasks. The novel $μ^2$Tokenizer, as an intermediate layer, integrates multi-modal features from the multiscale visual tokenizer and the text tokenizer, then enhances report generation quality through direct preference optimization (DPO), guided by GREEN-RedLlama. Experimental results on four large CT image-report medical datasets demonstrate that our method outperforms existing approaches, highlighting the potential of our fine-tuned $μ^2$LLMs on limited data for RRG tasks. At the same time, for prompt engineering, we introduce a five-stage, LLM-driven pipeline that converts routine CT reports into paired visual-question-answer triples and citation-linked reasoning narratives, creating a scalable, high-quality supervisory corpus for explainable multimodal radiology LLM. All code, datasets, and models will be publicly available in our official repository. https://github.com/Siyou-Li/u2Tokenizer
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。