用最新新闻自动构建测试集,评估大模型实时搜索能力。
LiveNewsBench: Evaluating LLM Web Search Capabilities with Freshly Curated News
- 从最新新闻自动生成需多步搜索的问题,检验模型真实检索能力。
- 包含人工验证样本,确保评估结果可靠;支持频繁更新数据。
- 适合研究智能搜索、评测大模型实时信息获取能力的学者使用。
具备代理式网络搜索能力的大语言模型在需要实时信息访问和复杂事实检索的任务中展现出巨大潜力,但评估此类系统仍具挑战性。我们提出 ench,一个严格且定期更新的基准,用于评估大模型的代理式网络搜索能力。该基准通过近期新闻文章自动生成新鲜的问答对,确保问题所需信息超出模型训练数据范围,从而清晰区分内部知识与搜索能力。基准包含需多跳搜索、页面访问和推理的难题,适合评估代理式搜索行为。自动化数据整理与问题生成流程支持频繁更新,并可构建大规模训练数据集,缓解研究社区中此类数据稀缺的问题。为保证评估可靠性,测试集包含人工验证样本。我们在 ench 上评估了多种系统,包括商业和开源大模型及基于大模型的网络搜索API。排行榜、数据集与代码已公开于 livenewsbench.com。
原文摘要 · Abstract (English)
Large Language Models (LLMs) with agentic web search capabilities show strong potential for tasks requiring real-time information access and complex fact retrieval, yet evaluating such systems remains challenging. We introduce \bench, a rigorous and regularly updated benchmark designed to assess the agentic web search abilities of LLMs. \bench automatically generates fresh question-answer pairs from recent news articles, ensuring that questions require information beyond an LLM's training data and enabling clear separation between internal knowledge and search capability. The benchmark features intentionally difficult questions requiring multi-hop search queries, page visits, and reasoning, making it well-suited for evaluating agentic search behavior. Our automated data curation and question generation pipeline enables frequent benchmark updates and supports construction of a large-scale training dataset for agentic web search models, addressing the scarcity of such data in the research community. To ensure reliable evaluation, we include a subset of human-verified samples in the test set. We evaluate a broad range of systems using \bench, including commercial and open-weight LLMs as well as LLM-based web search APIs. The leaderboard, datasets, and code are publicly available at livenewsbench.com.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。