arXiv:2602.22143cs.CV2026-02被引 2

将医学报告转为结构化三元组,提升视觉语言预训练效果

MedTri: A Platform for Structured Medical Report Normalization to Enhance Vision-Language Pretraining

  • 将自由文本报告统一转化为解剖部位:影像描述+诊断类别三元组
  • 在X光和CT数据上均显著优于原始报告与现有归一化方法
  • 支持知识增强与反事实监督等模块化增广策略,适合医疗AI研究者

医学视觉语言预训练日益依赖医学报告作为大规模监督信号,但原始报告常存在风格差异大、长度不一及大量与图像无关内容的问题。尽管文本归一化常被用作预处理,其设计原则与对预训练的影响尚未得到充分系统研究。本文提出MedTri,一个可部署的医学视觉语言预训练归一化框架,将自由文本报告转换为统一的[解剖部位:影像描述+诊断类别]三元组。该结构化、以解剖为基础的归一化保留了关键形态与空间信息,同时去除风格噪声与无关内容,在大规模图像-文本监督下提供一致且图像相关的语义信号。在涵盖X射线和计算机断层扫描(CT)模态的多个数据集上,我们证明结构化、以解剖为基础的文本归一化是提升医学视觉语言预训练质量的关键因素,相比原始报告与现有归一化基线均有稳定提升。此外,该归一化可轻松支持模块化文本级增强策略,如知识丰富与解剖基础反事实监督,显著提升模型鲁棒性与泛化能力,无需改变核心归一化流程。结果表明,结构化文本归一化是医学视觉语言学习中关键且通用的预处理组件,MedTri则为此提供了完整平台。代码与数据将在https://github.com/Arturia-Pendragon-Iris/MedTri公开。

原文摘要 · Abstract (English)

Medical vision-language pretraining increasingly relies on medical reports as large-scale supervisory signals; however, raw reports often exhibit substantial stylistic heterogeneity, variable length, and a considerable amount of image-irrelevant content. Although text normalization is frequently adopted as a preprocessing step in prior work, its design principles and empirical impact on vision-language pretraining remain insufficiently and systematically examined. In this study, we present MedTri, a deployable normalization framework for medical vision-language pretraining that converts free-text reports into a unified [Anatomical Entity: Radiologic Description + Diagnosis Category] triplet. This structured, anatomy-grounded normalization preserves essential morphological and spatial information while removing stylistic noise and image-irrelevant content, providing consistent and image-grounded textual supervision at scale. Across multiple datasets spanning both X-ray and computed tomography (CT) modalities, we demonstrate that structured, anatomy-grounded text normalization is an important factor in medical vision-language pretraining quality, yielding consistent improvements over raw reports and existing normalization baselines. In addition, we illustrate how this normalization can easily support modular text-level augmentation strategies, including knowledge enrichment and anatomy-grounded counterfactual supervision, which provide complementary gains in robustness and generalization without altering the core normalization process. Together, our results position structured text normalization as a critical and generalizable preprocessing component for medical vision-language learning, while MedTri provides this normalization platform. Code and data will be released at https://github.com/Arturia-Pendragon-Iris/MedTri.

医学AI视觉语言文本归一化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。