提出新评估标准,区分信息相关性与实际有用性。
UsefulBench: Towards Decision-Useful Information as a Target for Information Retrieval
- 构建专业标注数据集,区分文本相关性与实用价值
- 大模型在实用信息检索上仍存明显能力缺口
- 适合研究精准信息检索与LLM推理能力的学者
传统信息检索关注文本与查询的相似性,但忽视其是否真正有用。例如回答‘巴黎是否比柏林大’时,提及巴黎在法国的文本虽相关却无用。本文提出 UsefulBench,由三位专业分析师标注文本在回答查询时的实用性,建立领域专用数据集。实验表明,传统相似性方法更契合相关性判断;大语言模型虽可缓解偏差,但在需专业知识的场景中仍表现不足。该数据集为面向实用性的信息检索系统提供了评测基准。
原文摘要 · Abstract (English)
Conventional information retrieval is concerned with identifying the relevance of texts for a given query. Yet, the conventional definition of relevance is dominated by aspects of similarity in texts, leaving unobserved whether the text is truly useful for addressing the query. For instance, when answering whether Paris is larger than Berlin, texts about Paris being in France are relevant (lexical/semantic similarity), but not useful. In this paper, we introduce UsefulBench, a domain-specific dataset curated by three professional analysts labeling whether a text is connected to a query (relevance) or holds practical value in responding to it (usefulness). We show that classic similarity-based information retrieval aligns more strongly with relevance. While LLM-based systems can counteract this bias, we find that domain-specific problems require a high degree of expertise, which current LLMs do not fully incorporate. We explore approaches to (partially) overcome this challenge. However, UsefulBench presents a dataset challenge for targeted information retrieval systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。