评测检索系统能否找到复杂议题的多元观点。
Open-World Evaluation for Retrieving Diverse Perspectives
- 用真实辩论数据构建评估基准,衡量检索多样性。
- 现有模型仅在40%问题中覆盖全部观点,表现有限。
- 适合关注信息偏见、舆论多样性研究者使用。
我们研究如何检索一组涵盖复杂争议性问题(如‘ChatGPT带来的危害是否大于益处?’)多元观点的文档。为此,我们构建了针对主观问题的检索多样性基准(BERDS),每个样本包含一个问题及其来自问卷和辩论网站的多种观点。在该数据集上,评估不同检索器与语料库组合的表现,以找出包含多样观点的文档集合。不同于传统检索任务依赖关键词匹配,我们采用基于语言模型的自动评价器判断每篇文档是否包含特定观点。实验评估了三种语料库(维基百科、网页快照、实时搜索获取的语料)的表现,发现现有检索器仅在40%的样本中能覆盖所有观点。进一步分析了查询扩展和多样性重排序的效果,并探讨了检索器的‘谄媚倾向’现象。
原文摘要 · Abstract (English)
We study retrieving a set of documents that covers various perspectives on a complex and contentious question (e.g., will ChatGPT do more harm than good?). We curate a Benchmark for Retrieval Diversity for Subjective questions (BERDS), where each example consists of a question and diverse perspectives associated with the question, sourced from survey questions and debate websites. On this data, retrievers paired with a corpus are evaluated to surface a document set that contains diverse perspectives. Our framing diverges from most retrieval tasks in that document relevancy cannot be decided by simple string matches to references. Instead, we build a language model-based automatic evaluator that decides whether each retrieved document contains a perspective. This allows us to evaluate the performance of three different types of corpus (Wikipedia, web snapshot, and corpus constructed on the fly with retrieved pages from the search engine) paired with retrievers. Retrieving diverse documents remains challenging, with the outputs from existing retrievers covering all perspectives on only 40% of the examples. We further study the effectiveness of query expansion and diversity-focused reranking approaches and analyze retriever sycophancy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。