测试发现多数检索模型在处理否定时表现接近随机,仅新式大模型重排器稍好。
Reproducing NevIR: Negation in Neural Information Retrieval
- 用原基准复现并扩展实验,评估最新检索模型对否定的处理能力。
- 新出现的列表级大模型重排器性能最好,但仍远低于人类水平。
- 不同否定数据集间迁移效果差,说明数据分布差异大,适合关注模型泛化者看。
否定是人类交流的核心特征,但对语言模型在信息检索中的表现仍是挑战。尽管现代神经信息检索系统高度依赖语言模型,对其否定处理能力的关注却很少。本研究复现并拓展了NevIR的研究成果,该基准揭示大多数检索模型在否定任务中表现相当于或低于随机排序。我们复现了原始实验,并评估了新近发展的最先进检索模型。结果表明,一类新兴的列表级大语言模型重排器表现最佳,但仍显著落后于人类水平。此外,我们利用专为排除性查询设计的ExcluIR基准数据集,评估否定理解的泛化能力。发现在一个数据集上微调并不能可靠提升另一数据集的表现,表明两者数据分布存在显著差异。进一步观察到,仅交叉编码器和列表级大模型重排器在两类否定任务中均表现出合理性能。
原文摘要 · Abstract (English)
Negation is a fundamental aspect of human communication, yet it remains a challenge for Language Models (LMs) in Information Retrieval (IR). Despite the heavy reliance of modern neural IR systems on LMs, little attention has been given to their handling of negation. In this study, we reproduce and extend the findings of NevIR, a benchmark study that revealed most IR models perform at or below the level of random ranking when dealing with negation. We replicate NevIR's original experiments and evaluate newly developed state-of-the-art IR models. Our findings show that a recently emerging category-listwise Large Language Model (LLM) re-rankers-outperforms other models but still underperforms human performance. Additionally, we leverage ExcluIR, a benchmark dataset designed for exclusionary queries with extensive negation, to assess the generalisability of negation understanding. Our findings suggest that fine-tuning on one dataset does not reliably improve performance on the other, indicating notable differences in their data distributions. Furthermore, we observe that only cross-encoders and listwise LLM re-rankers achieve reasonable performance across both negation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。