arXiv:2609.08174cs.AI2026-09

测试密集检索在生物医学结构化约束下的表现,发现其对复杂关系理解不足。

OntologyBench: Can Dense Retrieval Satisfy Structured Biomedical Constraints?

论文配图:OntologyBench: Can Dense Retrieval Satisfy Structured Biomedical Constraints?
图 1 · 摘自论文原文
  • 构建包含47万训练、12.5万验证对的生物医学检索基准
  • 关系与组合任务准确率显著低于概念定位任务
  • 微调可提升性能,但现有重排序和大模型方法效果有限

我们提出OntologyBench,一个分层生物医学检索基准,涵盖471,854条训练和125,744条评估的查询-文档相关性对,覆盖概念定位、关系检索与组合表型检索三类任务。尽管这些任务可用本体感知的参考方法解决,但嵌入模型在关系与组合任务上的表现普遍低于概念定位任务。在本体导出的监督下进行微调,可提升多个关系与组合任务的表现;而评估的重排序与LLM候选评分方法则未带来明显端到端改进。错误常表现为疾病仅匹配部分表型证据。结果表明,现有嵌入与重排序配置无法可靠还原所选本体关系与表型组合中编码的兼容性,促使开发更融合学习表示与结构化生物医学知识的检索系统。

原文摘要 · Abstract (English)

We introduce OntologyBench, a tiered biomedical retrieval benchmark comprising 471,854 training and 125,744 evaluation query-document relevance pairs across concept grounding, relational retrieval, and compositional phenotype-based retrieval. Although these tasks can be tractable using ontology-aware reference methods, across task tiers, embedding performance is generally lower on relational and compositional tasks than on concept-grounding tasks. Fine-tuning on ontology-derived supervision improves performance on several relational and compositional tasks, whereas the evaluated reranking and LLM-based candidate-scoring methods provide little or no end-to-end improvement. Errors frequently reflect diseases matching only subsets of the phenotype evidence. These findings indicate that the evaluated embedding and reranking configurations do not reliably recover the compatibility encoded by the selected ontology relations and phenotype combinations and motivate retrieval systems that better integrate learned representations with structured biomedical knowledge.

信息检索本体生物医学嵌入模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。