基于知识的多样化重排序提升跨源问答检索效果
Knowledge-Aware Diverse Reranking for Cross-Source Question Answering
- 融合知识信息实现多样化重排序,增强文档相关性判断
- 在1500万文档子集上取得竞赛第一名,表现优异
- 适合需要高精度跨源问答的应用场景
本文介绍了团队Marikarp在SIGIR 2025 LiveRAG竞赛中的解决方案。竞赛评估集由DataMorgana从互联网语料自动生成,涵盖广泛的目标主题、问题类型、表达方式、受众群体及知识组织形式。该数据集对从FineWeb语料库1500万文档子集中检索与问题相关的支持文档提供了公平评测。我们提出的知识感知多样化重排序RAG流程在竞赛中获得第一名。
原文摘要 · Abstract (English)
This paper presents Team Marikarp's solution for the SIGIR 2025 LiveRAG competition. The competition's evaluation set, automatically generated by DataMorgana from internet corpora, encompassed a wide range of target topics, question types, question formulations, audience types, and knowledge organization methods. It offered a fair evaluation of retrieving question-relevant supporting documents from a 15M documents subset of the FineWeb corpus. Our proposed knowledge-aware diverse reranking RAG pipeline achieved first place in the competition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。