arXiv:2607.14581cs.AIcs.CV2026-07

用多大模型协作生成脑肿瘤MRI报告,提升诊断准确性。

Multi-LLM Collaborative MRI Report Generation for Visual Instruction Tuning in Brain Oncology

论文配图:Multi-LLM Collaborative MRI Report Generation for Visual Instruction Tuning in Brain Oncology
图 1 · 摘自论文原文
  • 多LLM协同生成并校验报告,保证准确清晰。
  • 构建首个3D MRI-文本数据集,支持脑瘤影像报告生成。
  • 新模型在报告生成和视觉问答上优于2D/3D方法。

大型语言模型(LLMs)及其扩展的视觉语言模型(VLMs)使得图像与文本结合的任务(如报告生成)更加容易。现有医学VLM主要针对2D图像(如胸部X光),难以拓展到3D成像,因缺乏配对的3D影像-文本数据。为此,我们提出一种新方法,利用胶质瘤和脑膜瘤患者的3D MRI扫描构建首个3D影像-文本数据集。采用多模型协作机制,多个LLM协同生成并验证报告,确保内容准确、表达清晰。基于该数据集,我们进一步构建了一个将MRI扫描转为标记(tokens)并与其文本指令对齐的VLM。该模型在报告生成和视觉问答任务中均优于其他2D及3D方法。本方法不仅提升了报告质量,还有助于脑瘤更精准的诊断与治疗。

原文摘要 · Abstract (English)

Recent advances in large language models (LLMs) and their extension to vision-language models (VLMs) have made it easier to combine text and images for tasks such as report generation. Existing VLMs in medicine typically focus on 2D images (chest X-rays), and their extension to 3D imaging has been difficult because of the lack of paired 3D imaging-text data. Thus, we introduce a new method for creating a 3D image-text dataset for brain oncology using 3D MRI scans of glioma and meningioma cases. We use a cooperative system in which several LLMs work together to generate and check reports, ensuring that they are accurate and clear. By leveraging the new 3D MRI-text dataset, we further build a VLM that converts MRI scans into tokens and aligns them with text instructions. Our VLM performed better in report generation and visual question answering tasks than other 2D and 3D methods. Our method not only improves the quality of reports but also helps with better diagnosis and treatment in brain oncology.

MRI报告多模型协作视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。