arXiv:2603.01690cs.CLcs.AI2026-03

用医学本体生成可解释的医疗文本嵌入,让模型决策像医生一样有依据。

QIME: Constructing Interpretable Medical Text Embeddings via Ontology-Grounded Questions

  • 基于医学本体生成每个维度对应一个临床问题的可解释嵌入
  • 无需训练即可构建嵌入,在相似度、聚类、检索任务上超越已有可解释方法
  • 适合需要透明决策过程的临床辅助系统,尤其关注可解释性需求

尽管密集型生物医学嵌入表现优异,但其黑箱特性限制了在临床决策中的应用。现有基于问题的可解释嵌入将文本表示为自然语言问题的二元回答,但通常依赖启发式或表层对比信号,忽视专业领域知识。本文提出QIME,一种基于本体的可解释医疗文本嵌入框架,其中每个维度对应一个临床上有意义的“是/否”问题。通过结合聚类特定的医学概念特征,QIME生成语义原子化的问题,捕捉生物医学文本的细微差异。此外,QIME支持免训练的嵌入构建策略,避免逐问题分类器训练,同时进一步提升性能。在生物医学语义相似性、聚类和检索基准测试中,QIME持续优于先前的可解释嵌入方法,并显著缩小与强黑箱生物医学编码器之间的差距,同时提供简洁且临床相关的解释。

原文摘要 · Abstract (English)

While dense biomedical embeddings achieve strong performance, their black-box nature limits their utility in clinical decision-making. Recent question-based interpretable embeddings represent text as binary answers to natural-language questions, but these approaches often rely on heuristic or surface-level contrastive signals and overlook specialized domain knowledge. We propose QIME, an ontology-grounded framework for constructing interpretable medical text embeddings in which each dimension corresponds to a clinically meaningful yes/no question. By conditioning on cluster-specific medical concept signatures, QIME generates semantically atomic questions that capture fine-grained distinctions in biomedical text. Furthermore, QIME supports a training-free embedding construction strategy that eliminates per-question classifier training while further improving performance. Experiments across biomedical semantic similarity, clustering, and retrieval benchmarks show that QIME consistently outperforms prior interpretable embedding methods and substantially narrows the gap to strong black-box biomedical encoders, while providing concise and clinically informative explanations.

可解释性医疗文本本体嵌入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。