arXiv:2603.08077cs.IR2026-03ACL

大模型推理可突破嵌入相似度局限,但现有数据集难验证其优势

Why Large Language Models can Secretly Outperform Embedding Similarity in Information Retrieval

  • 用大模型推理判断相关性,突破传统相似度匹配的短视局限
  • 在TREC-DL 2019上,带推理的LLM-RJS未显优,因标注本身有短视偏差
  • 误判多源于人工标注缺陷,说明需新评估方式验证模型真实潜力

随着大语言模型(LLM)的兴起,信息检索中可通过语言理解与推理直接判断相关性,而非依赖嵌入相似度。我们指出,相似度是相关性的短视解释,基于大模型的相关性判断系统(LLM-RJS,含推理)有望超越神经嵌入检索系统(NERS),因其能克服这一局限。我们在TREC-DL 2019片段检索数据集上对比多种LLM-RJS与NERS,未观察到明显提升。随后分析推理的影响,发现人工标注同样存在短视问题,且推理型LLM-RJS中的假阳性主要源于标注错误。结论:LLM-RJS具备解决NERS短视问题的能力,但标准标注相关性数据集无法有效评估其真实性能。

原文摘要 · Abstract (English)

With the emergence of Large Language Models (LLMs), new methods in Information Retrieval are available in which relevance is estimated directly through language understanding and reasoning, instead of embedding similarity. We argue that similarity is a short-sighted interpretation of relevance, and that LLM-Based Relevance Judgment Systems (LLM-RJS) (with reasoning) have potential to outperform Neural Embedding Retrieval Systems (NERS) by overcoming this limitation. Using the TREC-DL 2019 passage retrieval dataset, we compare various LLM-RJS with NERS, but observe no noticeable improvement. Subsequently, we analyze the impact of reasoning by comparing LLM-RJS with and without reasoning. We find that human annotations also suffer from short-sightedness, and that false-positives in the reasoning LLM-RJS are primarily mistakes in annotations due to short-sightedness. We conclude that LLM-RJS do have the ability to address the short-sightedness limitation in NERS, but that this cannot be evaluated with standard annotated relevance datasets.

信息检索大模型推理标注偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。