arXiv:2605.06235cs.IRcs.AI2026-05被引 4

提出新基准,测试模型发现隐含模式文档的能力

OBLIQ-Bench: Exposing Overlooked Bottlenecks in Modern Retrievers with Latent and Implicit Queries

  • 设计五类隐性查询任务,针对长尾真实数据集
  • 发现检索系统常漏掉相关文档,而验证模型能识别
  • 适合研究检索架构与隐式信号捕捉的学者

检索基准正趋于饱和,但我们认为高效搜索远未解决。本文提出一类称为‘斜向查询’(oblique)的查询类型,即寻找体现潜在模式的文档,如表达隐含立场的推文、展示特定故障模式的聊天记录、或匹配抽象场景的转录文本。我们分析了三种导致斜向性出现的机制,并构建了 OBLIQ-Bench——一个基于真实长尾语料库的五个斜向检索任务集合。该基准揭示了检索与验证之间的被忽视不对称:当相关文档被召回时,推理型大模型能可靠识别其隐含相关性,但即使先进的检索流水线也难以在首轮召回大多数相关文档。我们希望 OBLIQ-Bench 能推动检索架构研究,使其更高效地捕捉大规模语料中的潜在模式与隐式信号。

原文摘要 · Abstract (English)

Retrieval benchmarks are increasingly saturating, but we argue that efficient search is far from a solved problem. We identify a class of queries we call oblique, which seek documents that instantiate a latent pattern, like finding all tweets that express an implicit stance, chat logs that demonstrate a particular failure mode, or transcripts that match an abstract scenario. We study three mechanisms through which obliqueness may arise and introduce OBLIQ-Bench, a suite of five oblique search problems over real long-tail corpora. OBLIQ-Bench exposes an overlooked asymmetry between retrieval and verification, where reasoning LLMs reliably recognize latent relevance whenever relevant documents are surfaced, but even sophisticated retrieval pipelines fail to surface most relevant documents in the first place. We hope that OBLIQ-Bench will drive research into retrieval architectures that efficiently capture latent patterns and implicit signals in large corpora.

信息检索隐式查询评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。