arXiv:2504.15929cs.CVcs.AI2025-04被引 11

通过细粒度病灶描述增强医学影像与文本对齐,提升肺部X光诊断模型效果。

Meta-Entity Driven Triplet Mining for Aligning Medical Vision-Language Models

  • 基于疾病类别和病灶属性(位置/大小/严重程度)构建元实体,指导三元组挖掘
  • 在公开数据集上实现更优的图像-文本检索与分类性能,提升关键任务指标
  • 适合需要精准病理描述理解的医疗AI研究者和临床辅助系统开发者

诊断影像依赖图像与放射科报告的联合解读,但数据量激增给医疗专家带来巨大压力,导致错误率上升和工作流程积压。医学视觉语言模型(med-VLMs)作为高效处理多模态影像数据的框架,尤其在胸部X光(CXR)评估中表现突出,其性能高度依赖图像与文本表征的对齐程度。现有对齐方法主要基于对比学习,侧重于区分疾病大类,却忽视了位置、大小、严重程度等细微病理特征的分离,导致表征不充分。本文提出MedTrim(元实体驱动的三元组挖掘),一种新型方法,通过融合疾病类别及形容词性、方向性病理描述的多模态三元组学习,提升图像-文本对齐。不同于常规方法仅分离大类,MedTrim利用结构化元实体信息保留同类内的细微差异。为此,我们引入基于本体的实体识别模块,从CXR报告中提取病灶特异性元实体——因公开数据集中病灶属性标注稀缺,此步骤至关重要。为优化三元组样本选择,我们设计了一种新颖的评分函数,综合疾病类别与形容词/方向描述,衡量样本间相似性。最后,提出多模态三元组对齐目标,显式强化共享详细病理特征的样本间的跨模态对齐。实验表明,相较于最先进对齐方法,MedTrim在下游检索与分类任务中均取得显著性能提升。

原文摘要 · Abstract (English)

Diagnostic imaging relies on interpreting both images and radiology reports, but the growing data volumes place significant pressure on medical experts, yielding increased errors and workflow backlogs. Medical vision-language models (med-VLMs) have emerged as a powerful framework to efficiently process multimodal imaging data, particularly in chest X-ray (CXR) evaluations, albeit their performance hinges on how well image and text representations are aligned. Existing alignment methods, predominantly based on contrastive learning, prioritize separation between disease classes over segregation of fine-grained pathology attributes like location, size or severity, leading to suboptimal representations. Here, we propose MedTrim (Meta-entity-driven Triplet mining), a novel method that enhances image-text alignment through multimodal triplet learning synergistically guided by disease class as well as adjectival and directional pathology descriptors. Unlike common alignment methods that separate broad disease classes, MedTrim leverages structured meta-entity information to preserve subtle but clinically significant intra-class variations. For this purpose, we first introduce an ontology-based entity recognition module that extracts pathology-specific meta-entities from CXR reports, as annotations on pathology attributes are rare in public datasets. For refined sample selection in triplet mining, we then introduce a novel score function that captures an aggregate measure of inter-sample similarity based on disease classes and adjectival/directional descriptors. Lastly, we introduce a multimodal triplet alignment objective for explicit within- and cross-modal alignment between samples sharing detailed pathology characteristics. Our demonstrations indicate that MedTrim improves performance in downstream retrieval and classification tasks compared to state-of-the-art alignment methods.

医学多模态视觉语言对齐影像诊断三元组学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。