评测智能代理在热门话题中多源信息整合能力。
iAgentBench: Benchmarking Sensemaking Capabilities of Information-Seeking Agents on High-Traffic Topics
- 基于真实热点话题构建需跨源整合的问答任务
- 单次检索无法解决复杂问题,需多步推理与证据融合
- 提供可追溯证据链,适合评估大模型综合分析能力
随着具备搜索功能的生成式问答系统兴起,用户越来越依赖能自动浏览、聚合并协调多个来源证据的工具。然而,许多常用问答基准仍可通过单一相关段落回答,难以衡量跨源理解、因果追踪和多维度矛盾化解等高级信息处理能力。我们提出iAgentBench,一个动态的开放域问答基准,聚焦真实信息寻求行为中的高阶需求。该基准从现实关注度信号中选取主题,依据常见用户意图模式生成需整合多源信息的问题。每个实例均附带可追溯的证据与可审计的中间产物,支持污染检测与故障细粒度诊断。在多个大语言模型上的实验表明,检索提升准确率,但仅靠检索无法可靠解答问题,凸显了评估证据使用而非仅证据获取的重要性。
原文摘要 · Abstract (English)
With the emergence of search-enabled generative QA systems, users are increasingly turning to tools that browse, aggregate, and reconcile evidence across multiple sources on their behalf. Yet many widely used QA benchmarks remain answerable by retrieving a single relevant passage, making them poorly suited for measuring cross-source sensemaking, such as integrating evidence, tracking causal links, and resolving dependencies across facets of a topic. We present iAgentBench, a dynamic ODQA benchmark that targets these higher-level information needs while keeping questions natural and grounded in realistic information-seeking behavior. iAgentBench draws seed topics from real-world attention signals and uses common user intent patterns to construct user-like questions whose answers require combining evidence from multiple sources, not just extracting a single snippet. Each instance is released with traceable evidence and auditable intermediate artifacts that support contamination checks and enable fine-grained diagnosis of failures in retrieval versus synthesis. Experiments across multiple LLMs show that retrieval improves accuracy, but retrieval alone does not reliably resolve these questions, underscoring the need to evaluate evidence use, not just evidence access.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。