arXiv:2605.14002cs.AI2026-05ACL

构建多语言政治人物事实发现基准,评估模型在长尾信息中的探索能力

PolitNuggets: Benchmarking Agentic Discovery of Long-Tail Political Facts

论文配图:PolitNuggets: Benchmarking Agentic Discovery of Long-Tail Political Facts
图 1 · 摘自论文原文
  • 构建400位全球政要的多语种政治传记,覆盖超1万条政治事实
  • 提出FactNet评估协议,量化发现、精度与效率,揭示现有系统短板
  • 诊断显示短上下文提取与多语言鲁棒性是关键性能因素

嵌入代理框架的大推理模型已将信息检索从静态长文本问答转变为开放式探索。然而真实应用要求模型从分散来源中发现并合成‘长尾’事实,这一能力仍缺乏有效评估。我们提出PolitNuggets,一个基于400位全球精英政治传记的多语言基准,涵盖超过10000条政治事实。通过优化的多代理系统实现标准化评估,并提出FactNet——一种基于证据的评分协议,用于衡量发现能力、细粒度准确性及效率。在不同模型与设置下,我们发现当前系统普遍难以处理细粒度细节,且效率差异显著。最后,借助基准诊断,我们将代理性能与底层模型能力关联,凸显短上下文提取、多语言鲁棒性及可靠工具使用的重要性。

原文摘要 · Abstract (English)

Large Reasoning Models (LRMs) embedded in agentic frameworks have transformed information retrieval from static, long context question answering into open-ended exploration. Yet real world use requires models to discover and synthesize "long-tail" facts from dispersed sources, a capability that remains under-evaluated. We introduce PolitNuggets, a multilingual benchmark for agentic information synthesis via constructing political biographies for 400 global elites, covering over 10000 political facts. We standardize evaluation with an optimized multi agent system and propose FactNet, an evidence conditional protocol that scores discovery, fine-grained accuracy, and efficiency. Across models and settings, we find that current systems often struggle with fine-grained details, and vary substantially in efficiency. Finally, using benchmark diagnostics, we relate agent performance to underlying model capabilities, highlighting the importance of short-context extraction, multilingual robustness, and reliable tool use.

信息发现多语言评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。