解决医学图文检索中的语义碎片化问题,提升跨模态匹配精度。
TriPAH: Imbalance-Aware Tri-Prompt Affinity Hashing for Cross-Modal Medical Retrieval

- 设计三重提示框架,融合医学本体与患者特征生成低噪声文本表征。
- 在三个公开数据集上均超越当前最优方法,尤其在长尾标签场景下提升显著。
- 适合医疗AI研发者、医学图像检索系统开发者使用。
在大数据医疗时代,高效的跨模态检索对循证诊断和大规模病例管理至关重要。跨模态医学哈希检索通过学习紧凑且语义对齐的二值编码,实现图像-文本高效搜索,并支持基于案例推理与决策支持等下游任务。然而,现有方法因临床语言噪声、长尾标签分布及脆弱的量化过程导致语义碎片化,削弱了对齐效果。本文提出TriPAH——一种三提示亲和哈希框架。该框架利用医学本体引导、基于标准化临床线索的患者级提示,生成低噪声文本表征以实现初始对齐;轻量级提示-令牌混合器在异构多任务目标下进行分层、多粒度对齐,输出可量化的特征,该目标联合了多正例对比对齐、不平衡感知分类与渐进式量化正则化;患者级一致性模块进一步稳定不同视图下的编码表现。在三个公开数据集上的大量实验表明,TriPAH显著优于现有先进方法。
原文摘要 · Abstract (English)
In the era of big medical data, efficient cross-modal retrieval is pivotal for evidence-based diagnosis and large-scale case management. Cross-modal medical hashing retrieval aims to enable efficient image-text search and support downstream tasks such as case-based reasoning and decision support by learning compact, semantically aligned binary codes. However, current methods suffer from semantic fragmentation due to noisy clinical language, long-tailed labels, and brittle quantization that weakens alignment. We propose TriPAH, a Tri-Prompt Affinity Hashing framework. TriPAH synthesizes ontology-grounded, patient-level prompts conditioned on normalized clinical cues to yield low-noise textual representations for initial alignment. A lightweight prompt-token mixer performs hierarchical, multi-granularity alignment and produces quantization-ready features under an asymmetric multi-task objective coupling multi-positive contrastive alignment, imbalance-aware classification, and progressive quantization regularization. A patient-level consistency module further stabilizes codes across complementary views. Extensive experiments on three public datasets demonstrate that TriPAH significantly outperforms state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。