让ColBERT模型的检索结果可诊断,用临床知识对齐嵌入空间。
Diagnosable ColBERT: Debugging Late-Interaction Retrieval Models Using a Learned Latent Space as Reference

- 用临床知识构建参考潜空间,对齐ColBERT的词元嵌入。
- 通过嵌入空间位置判断模型是否稳定理解临床概念。
- 无需大量诊断查询,即可定位模型错误并指导数据优化。
可靠的生物医学和临床检索不仅需要强大的排序性能,还需实用的方法来发现系统性模型缺陷,并整理训练所需证据以修正错误。晚交互模型如ColBERT通过暴露文档与查询词元间的可解释交互得分,提供了一种初步解决方案。然而这种可解释性是浅层的:它仅解释特定文档-查询配对得分,无法揭示模型是否以稳定、可复用且上下文敏感的方式学习了临床概念。因此,这些得分难以支持对误解的诊断、异常远距离生物医学概念的识别,或判断需补充哪些数据或反馈。本文提出Diagnosable ColBERT框架,将ColBERT的词元嵌入对齐至基于临床知识和专家提供的概念相似性约束构建的参考潜空间。该对齐使文档编码成为模型似乎理解内容的可检查证据,从而实现更直接的错误诊断和更严谨的数据整理,无需依赖大规模诊断查询集。
原文摘要 · Abstract (English)
Reliable biomedical and clinical retrieval requires more than strong ranking performance: it requires a practical way to find systematic model failures and curate the training evidence needed to correct them. Late-interaction models such as ColBERT provide a first solution thanks to the interpretable token-level interaction scores they expose between document and query tokens. Yet this interpretability is shallow: it explains a particular document--query pairwise score, but does not reveal whether the model has learned a clinical concept in a stable, reusable, and context-sensitive way across diverse expressions. As a result, these scores provide limited support for diagnosing misunderstandings, identifying irreasonably distant biomedical concepts, or deciding what additional data or feedback is needed to address this. In this short position paper, we propose Diagnosable ColBERT, a framework that aligns ColBERT token embeddings to a reference latent space grounded in clinical knowledge and expert-provided conceptual similarity constraints. This alignment turns document encodings into inspectable evidence of what the model appears to understand, enabling more direct error diagnosis and more principled data curation without relying on large batteries of diagnostic queries.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。