让AI理解病理图文混合查询,提升医学检索可解释性。
Histopathology Multi-modal Embedding for Pathology Composed Retrieval
- 提出HOMIE框架,将通用多模态模型改造为病理检索专家。
- 在新基准上,20亿参数模型超越70亿参数专用模型。
- 适合临床辅助诊断、可解释AI研究者使用。
为克服预测型AI的黑箱问题和生成模型的幻觉风险,基于检索的模型提供了可解释、以证据为基础的病理临床工作流方案。然而,真实临床查询天然包含多模态混合(如病理图像与文本)。现有双编码器存在架构不匹配问题,缺乏融合此类复合查询的机制。为此,我们正式提出病理复合检索(Pathology Composed Retrieval, PCR)任务。尽管多模态大语言模型(MLLMs)具备深度融合能力,但直接应用存在任务错配与领域错配。为此,我们提出HOMIE——一种模型无关的适配框架,可将任意生成式MLLM转化为专业病理检索模型。在新提出的PCR基准上,一个轻量级20亿参数的HOMIE变体,在复合检索任务上显著优于现有范式,超越专用70亿参数病理MLLM和双编码器,同时保持对传统简单检索的良好性能。
原文摘要 · Abstract (English)
To overcome the black-box nature of predictive AI and the hallucination risks of generative models, retrieval-based models offer an interpretable, evidence-based paradigm for pathology clinical workflow. However, real-world clinical queries are inherently interleaved (e.g., pathology images and text). Current dual-encoders suffer from an \textbf{Architectural Mismatch}, lacking the mechanism to fuse such composed queries. To address this, we formalize the task of Pathology Composed Retrieval (PCR). While Multimodal Large Language Models (MLLMs) offer deep-fusion capabilities, directly applying them exposes a \textbf{Task Mismatch} and a \textbf{Domain Mismatch}. To resolve these challenges, we propose HOMIE, a model-agnostic adaptation framework that transforms any generative MLLM into a specialized pathology retrieval expert. Evaluated on our newly introduced PCR Benchmark, a lightweight 2B-parameter HOMIE variant substantially outperforms existing paradigms, surpassing specialized 7B pathology MLLMs and dual-encoders by large margins on composed retrieval, while maintaining strong performance on traditional simple retrieval. The project page is available at https://qfchou.github.io/HOMIE_page/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。