实体检索效果受限于信号覆盖不足,而非模型能力。
Entities as Retrieval Signals: A Systematic Study of Coverage, Supervision, and Evaluation in Entity-Oriented Ranking
- 区分概念相关性与可观察相关性,揭示实体信号瓶颈
- 实体信号仅覆盖19.7%相关文档,难以兼顾覆盖率与区分度
- 现有评估混淆条件与开放世界场景,误导性能判断
实体导向检索假设相关文档包含与查询相关的实体,但现有评估结果矛盾。在TREC Robust04上,我们评估了六种神经重排序器和437种无监督配置,均以BM25为基线。在443个系统中,开放世界评估下无一系统在完整候选集上将MAP提升超过0.05,尽管在实体受限设置下有显著增益。最佳配置达到官方Robust04最优系统水平,优于多数神经重排序器,说明模型架构非瓶颈。真正限制在于实体通道:即使理想选择,实体信号也仅覆盖19.7%的相关文档,且无方法能同时实现高覆盖与强区分。根源在于区分概念实体相关性(CER)与可观察实体相关性(OER)——所有监督策略基于CER,忽略链接环境,导致语义合理但缺乏区分力的信号。改进监督无法恢复开放世界性能:更强信号反而降低覆盖率。条件评估关注利用实体证据,开放世界评估关注真实链接下的检索效果,二者常被混淆。进展需具备实体层面区分力的数据集及同时报告覆盖与效果的评估。否则,条件增益不等于开放世界有效,开放世界失败也不否定实体模型价值。
原文摘要 · Abstract (English)
Entity-oriented retrieval assumes that relevant documents exhibit query-relevant entities, yet evaluations report conflicting results. We show this inconsistency stems not from model failure, but from evaluation. On TREC Robust04, we evaluate six neural rerankers and 437 unsupervised configurations against BM25. Across 443 systems, none improves MAP by more than 0.05 under open-world evaluation over the full candidate set, despite strong gains under entity-restricted settings. The best configuration matches the official Robust04 best system and outperforms most neural rerankers, indicating that the architecture is not the limiting factor. Instead, the bottleneck is the entity channel: even under idealized selection, entity signals cover only 19.7\% of relevant documents, and no method achieves both high coverage and discrimination. We explain this via a distinction between Conceptual Entity Relevance (CER) -- semantic relatedness -- and Observable Entity Relevance (OER) -- corpus-grounded discriminativeness under a given linker. All supervision strategies operate at the CER level and ignore the linking environment, leading to signals that are semantically valid but not discriminative. Improving supervision therefore does not recover open-world performance: stronger signals reduce coverage without improving effectiveness. Conditional and open-world evaluation answer different questions: exploiting entity evidence versus improving retrieval under realistic linking, but are often conflated. Progress requires datasets with entity-level discriminativeness and evaluation that reports both coverage and effectiveness. Until then, conditional gains do not imply open-world effectiveness, and open-world failures do not invalidate entity-based models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。