arXiv:2505.17162cs.IRcs.AI2025-05被引 4

DailyQA动态评测大模型对实时网络信息的处理能力

DailyQA: A Benchmark to Evaluate Web Retrieval Augmented LLMs Based on Capturing Real-World Changes

  • 每周自动更新问题与答案,基于维基百科修订日志构建
  • 重排序检索结果显著提升模型处理时效信息的能力
  • 适合评估大模型在快速变化事实上的真实表现

我们提出 DailyQA,一个可自动更新的动态数据集,每周更新问题并包含任意日期的答案。DailyQA 利用维基百科修订日志实现从数据筛选、问题生成、质量检查、答案抽取到查询分类的全自动流程。该基准要求大语言模型(LLMs)处理涉及快速变化的事实性信息,并覆盖多个领域。我们使用不同 RAG 流水线和网络搜索增强方法评估多个开源与闭源 LLMs。结果显示,重排序检索结果对处理时效性信息至关重要。实验表明,当前大模型在处理频繁更新信息方面仍面临显著挑战,说明 DailyQA 能为 LLM 与 RAG 系统的发展方向提供重要洞察。

原文摘要 · Abstract (English)

We propose DailyQA, an automatically updated dynamic dataset that updates questions weekly and contains answers to questions on any given date. DailyQA utilizes daily updates from Wikipedia revision logs to implement a fully automated pipeline of data filtering, query generation synthesis, quality checking, answer extraction, and query classification. The benchmark requires large language models (LLMs) to process and answer questions involving fast-changing factual data and covering multiple domains. We evaluate several open-source and closed-source LLMs using different RAG pipelines with web search augmentation. We compare the ability of different models to process time-sensitive web information and find that rerank of web retrieval results is critical. Our results indicate that LLMs still face significant challenges in handling frequently updated information, suggesting that DailyQA benchmarking provides valuable insights into the direction of progress for LLMs and RAG systems.

大模型评测RAG动态数据实时问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。