arXiv:2512.18679cs.CVcs.CL2025-12中稿 · WACV 2026

用临床报告对齐脑部MRI多视角表示,提升病灶检测能力

brat: Aligned Multi-View Embeddings for Brain MRI Analysis

  • 基于报告与MRI配对数据,设计跨模态对齐的多视图表示学习框架
  • 在8万例3D脑MRI上训练,性能显著优于现有方法
  • 适合医学影像分析、多模态模型研究者使用

我们提出brat(brain report alignment transformer),一种基于临床报告配对的脑部磁共振成像(MRI)多视图表征学习框架。脑MRI因存在大量、高度多样且常为局部细微异常,具有独特挑战。为此,我们构建了一个规模达现有数据集10倍的脑部MRI数据集,包含约8万个3D扫描及对应的放射科报告,并提出受文档检索启发的多视图预训练方法。通过隐式查询-特征匹配机制,结合质量-多样性思想,实现MRI多视图嵌入与报告句子中临床特征的对齐。我们在多个视觉-语言和视觉任务上评估该方法,表现出显著性能提升。brat基础模型已公开发布。

原文摘要 · Abstract (English)

We present brat (brain report alignment transformer), a multi-view representation learning framework for brain magnetic resonance imaging (MRI) trained on MRIs paired with clinical reports. Brain MRIs present unique challenges due to the presence of numerous, highly varied, and often subtle abnormalities that are localized to a few slices within a 3D volume. To address these challenges, we introduce a brain MRI dataset $10\times$ larger than existing ones, containing approximately 80,000 3D scans with corresponding radiology reports, and propose a multi-view pre-training approach inspired by advances in document retrieval. We develop an implicit query-feature matching mechanism and adopt concepts from quality-diversity to obtain multi-view embeddings of MRIs that are aligned with the clinical features given by report sentences. We evaluate our approach across multiple vision-language and vision tasks, demonstrating substantial performance improvements. The brat foundation models are publicly released.

脑部MRI多模态学习医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。