arXiv:2607.20498cs.AI2026-07KDD

构建学术知识图谱信息检索的全流程评测基准,挑战大模型多步调用API的能力。

AISE-Bench: A Full-Cycle Curated Benchmark for Information Seeking on Academic Knowledge Graphs

论文配图:AISE-Bench: A Full-Cycle Curated Benchmark for Information Seeking on Academic Knowledge Graphs
图 1 · 摘自论文原文
  • 设计真实用户意图的全周期任务,支持复杂API规划与参数填充
  • 14个模型平均表现中等,最强模型在API规划和执行上仍常出错
  • 适合评估大模型在科研场景下的多步推理与可追溯决策能力

大语言模型结合工具正成为能完成复杂长程任务的自主智能体。现有学术图谱信息检索评测多依赖合成模板、简化解空间或单一论文相关任务,未能覆盖真实用户意图、多步API规划、丰富参数填写、带参考来源的准确回答及全过程评估等关键挑战。我们提出AISE-Bench,一个基于真实场景的学术知识图谱信息检索全周期标注基准。该基准包含1,133组问答对,涵盖查询分类、完整API执行轨迹、验证参数及带引用链接的源文档答案。为保证标注质量,我们设计定制化智能体工作流,支持注释者高效规划、执行与修正复杂API流程。建立涵盖答案质量、参考依据、API规划正确性与执行成功率的综合评估协议。在14种方法中,即使最强模型(PLAY2PROMPT + Gemini-3-Pro)也仅达中等水平,且在API规划与执行上频繁失败。AISE-Bench为量化评估与提升多步调用模型的步骤正确性、有根据总结与可追踪推理能力提供了新基准。代码与数据已公开于https://aise-bench.github.io/。

原文摘要 · Abstract (English)

Large language models (LLMs) augmented with tools are emerging as autonomous agents capable of using Web engine, APIs, and code to solve complex, long-horizon tasks. Current tool-using benchmarks for information seeking on academic graphs rely on synthetic templates, simplified solution spaces, or narrow tasks such as paper-centric tasks, leaving key challenges underexplored - realistic user intent, complex multi-step API planning, rich parameter filling for APIs, grounded answers with references, and comprehensive evaluation of both the process and the outcome. We introduce AISE-Bench, a real-world, full-cycle annotated benchmark for information seeking on academic knowledge graphs. AISE-Bench release contains 1,133 QA pairs, including query taxonomies, full API execution trajectories, validated parameters, and source-grounded answers with reference links. To support high-quality annotation, we design a customized agent workflow to enable annotators to plan, execute, and revise complex API workflows efficiently. We develop a comprehensive evaluation protocol measuring answer quality, reference grounding, API-planning correctness, and execution success. Among the 14 evaluated methods, even the strongest model (PLAY2PROMPT with Gemini-3-Pro) achieves only moderate performance and often struggles with API planning and execution. AISE-Bench establishes a challenging new testbed for quantitatively evaluating and improving the stepwise correctness, grounded summarization, and traceable reasoning of multi-step API-using LLM agents. Our code and data are available at https://aise-bench.github.io/.

知识图谱信息检索大模型评测API规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。