arXiv:2507.04942cs.CLcs.IR2025-07

70支团队在2小时内挑战10亿条数据的RAG问答,比拼检索与提示策略。

SIGIR 2025 -- LiveRAG Challenge Report

  • 使用固定语料库和开源大模型,对比不同检索与提示方法。
  • 500个未见问题在两小时内完成回答,自动评分+人工复核。
  • 来自27国的顶尖团队角逐,成果将在SIGIR大会揭晓。

2025年3月至5月举行的SIGIR 2025 LiveRAG挑战赛,为推进检索增强生成(RAG)技术提供了竞技平台。来自学术界和产业界的参赛者需基于固定语料库(Fineweb-10BT)和通用开源大模型(Falcon3-10B-Instruct),开发一个基于RAG的问答系统,以促进对检索与提示策略的挑战性对比。在直播挑战日,来自27个国家的70支队伍在严格两小时时限内,回答了500个未见问题,并提供支持信息。评估分为两个阶段:首先采用LLM-as-a-judge进行自动化评分,计算正确性与忠实性;随后对排名靠前的提交结果进行人工评审。最终入围名单于2025年6月12日公布,奖项在意大利帕多瓦举行的LiveRAG研讨会期间颁发。

原文摘要 · Abstract (English)

The LiveRAG Challenge at SIGIR 2025, held between March and May 2025, provided a competitive platform for advancing Retrieval-Augmented Generation (RAG) technologies. Participants from academia and industry were invited to develop a RAG-based question-answering system using a fixed corpus (Fineweb-10BT) and a common open-source LLM (Falcon3-10B-Instruct). The goal was to facilitate challenging comparisons of retrieval and prompting strategies. During the Live Challenge Day, 70 teams from 27 different countries provided answers and supportive information to 500 unseen questions within a strict two-hour time window. Evaluation was conducted in two stages: first an automated LLM-as-a-judge approach was used to compute correctness and faithfulness score, then a manual review of top ranked submissions was conducted. The finalists were announced on June 12, 2025, with prizes awarded during the LiveRAG Workshop at SIGIR 2025 in Padua, Italy.

RAG问答系统大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。