通过图文联合对齐,提升皮肤癌病例检索精度。
Composed Vision-Language Retrieval for Skin Cancer Case Search via Joint Alignment of Global and Local Representations
- 构建分层图文表示,融合全局语义与局部特征。
- 在Derm7pt数据集上优于现有方法,显著提升检索准确率。
- 适合临床辅助诊断与医学影像知识库建设。
医学图像检索旨在识别临床相关的病变病例,以支持诊断决策、教学和质量控制。实际检索中,查询常包含参考病变图像与文本描述(如皮肤镜特征)。本文研究皮肤癌的组合视觉-语言检索,每个查询由图像-文本对构成,数据库包含经活检确认的多类疾病病例。提出一种基于Transformer的框架,学习分层的组合查询表示,并实现查询与候选图像之间的全局-局部联合对齐。局部对齐通过多重空间注意力掩码聚合判别性区域,全局对齐提供整体语义监督。最终相似度采用凸型、领域知情加权计算,突出临床关键局部证据,同时保持全局一致性。在公开数据集Derm7pt上的实验表明,该方法持续优于当前最优方法。所提框架可高效访问相关病历记录,支持临床实用部署。
原文摘要 · Abstract (English)
Medical image retrieval aims to identify clinically relevant lesion cases to support diagnostic decision making, education, and quality control. In practice, retrieval queries often combine a reference lesion image with textual descriptors such as dermoscopic features. We study composed vision-language retrieval for skin cancer, where each query consists of an image to text pair and the database contains biopsy-confirmed, multi-class disease cases. We propose a transformer based framework that learns hierarchical composed query representations and performs joint global-local alignment between queries and candidate images. Local alignment aggregates discriminative regions via multiple spatial attention masks, while global alignment provides holistic semantic supervision. The final similarity is computed through a convex, domain-informed weighting that emphasizes clinically salient local evidence while preserving global consistency. Experiments on the public Derm7pt dataset demonstrate consistent improvements over state-of-the-art methods. The proposed framework enables efficient access to relevant medical records and supports practical clinical deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。