arXiv:2606.04911cs.CVcs.CL2026-06

构建乳腺癌全流程多模态模型,提升临床推理能力

BreastGPT: A Multimodal Large Language Model for the Full Spectrum of Breast Cancer Clinical Routine

论文配图:BreastGPT: A Multimodal Large Language Model for the Full Spectrum of Breast Cancer Clinical Routine
图 1 · 摘自论文原文
  • 基于186万指令对构建全流程乳腺影像数据集
  • 在多阶段任务中实现75.7%闭合题准确率与89.9%开放题得分
  • 适合临床辅助决策系统研发者使用

乳腺癌是女性癌症致死的主要原因。其临床管理涉及筛查、诊断和治疗规划全流程,各阶段涵盖不同影像模态、任务目标与推理模式。然而,受限于数据稀缺与模型泛化能力,现有医学多模态大模型通常仅在单一模态或窄范围任务上评估,难以支持全流程临床推理。本文首次提出 extbf{BreastStage},一个与临床流程对齐的乳腺影像指令语料库,包含来自5种影像模态和136个任务模板的17个子数据集,共计186万条指令跟随样本。其预留测试集 extbf{BreastStage-Bench} 可全面评估跨乳腺癌全周期的多模态推理能力。在此基础上,我们提出 extbf{BreastGPT},一种统一的多模态大语言模型,采用双分支视觉编码器与概念保持的令牌压缩机制,弥合标准放射科影像与千兆像素病理图像间的尺度差距。在 extbf{BreastStage-Bench} 上,BreastGPT 实现75.66%闭合题准确率与89.92%开放题得分,优于通用及医学专用多模态大模型,在各临床阶段与任务格式下均表现更优。结果表明,流程对齐数据与跨尺度视觉建模对构建临床可信的医学多模态大模型至关重要。所有数据、代码与模型检查点已公开于 https://yangyy-liu.github.io/BreastGPT.io。

原文摘要 · Abstract (English)

Breast cancer remains a leading cause of cancer-related mortality among women. Its clinical management requires multimodal reasoning across a clinical workflow that spans \textit{screening}, \textit{diagnosis} and \textit{treatment planning}, where each stage involves distinct imaging modalities, task objectives, and reasoning patterns. However, constrained by data scarcity and model versatility, existing medical MLLMs are typically evaluated on isolated modalities or narrow task families, limiting their ability to support workflow-level clinical reasoning. In this work, we first introduce \textbf{BreastStage}, a workflow-aligned breast imaging instruction corpus comprising 1.86M instruction-following pairs curated from 17 sub-datasets across 5 imaging modalities and 136 task templates. Its held-out split, \textbf{BreastStage-Bench}, provides a comprehensive benchmark for evaluating multimodal reasoning across the breast cancer care continuum. Building on this corpus, we propose \textbf{BreastGPT}, a unified MLLM equipped with a dual-branch visual encoder and concept-preserving token compression to bridge the scale gap between standard radiology and gigapixel pathology. On BreastStage-Bench, BreastGPT achieves 75.66\% closed-ended accuracy and 89.92\% open-ended score, outperforming both general-purpose and medical-specific MLLMs across clinical stages and task formats. These results suggest that workflow-aligned data and cross-scale visual modeling are critical for clinically grounded medical MLLMs. All data, code, and model checkpoints are released at https://yangyy-liu.github.io/BreastGPT.io.

乳腺癌多模态大模型临床推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。