剖析ngram生成检索的失败原因,提供可交互诊断工具。
Understanding and Debugging Failures in N-Gram-Based Generative Retrieval

- 构建生成检索失败模式分类体系,聚焦ngram方法缺陷。
- 发现文档标识符歧义、多样性低及关键标识符过载等共性问题。
- 推出可视化工具,直观定位生成检索中的错误环节。
生成式检索(GR)是一种新兴的信息检索范式,依赖于日益强大的语言模型直接生成相关文档的标识符。尽管该系统具有独特优势,但也引入了特有的失败机制。本文有三项贡献:(1) 基于现有文献提出一种GR失败模式的分类体系;(2) 实证分析ngram-based方法(特别是SEAL和MINDER)中的失败现象,发现常见问题包括文档标识符歧义、标识符多样性低,以及特定标识符对最终排序产生不成比例的影响;(3) 提出一款新的基于Web的工具,帮助信息检索社区分析生成的ngram及其对最终排序的贡献,提供直观界面以识别GR方法的错误位置。
原文摘要 · Abstract (English)
Generative Retrieval (GR) is an emerging Information Retrieval (IR) paradigm that is motivated by increasingly capable language models. In GR, a model directly generates identifiers for relevant documents. While these systems offer unique advantages, they also introduce distinct failure mechanisms. We explore these failure modes in three contributions: (1) We present a taxonomy of GR failure modes based on GR literature. (2) We empirically investigate failure in a subset of GR: ngram-based methods, more specifically, SEAL and MINDER. Our analysis reveals common issues, such as ambiguous docids, low identifier diversity, and the disproportionate impact of specific identifiers. (3) We introduce a new web-based tool that helps the IR community analyze generated ngrams and their respective contribution to the final ranking, providing an intuitive interface to identify where such GR methods go wrong.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。