首个评估大模型搜索代理认知能力的基准,关注其证据推理与自我校准。
Do LLM Agents Know How to Ground, Recover, and Assess? A Benchmark for Epistemic Competence in Information-Seeking Agents
- 通过步骤级分析,评估代理是否基于真实证据推理。
- 发现多数代理无法有效从低质结果中恢复或判断证据是否充分。
- 适合研究智能体可信度、可解释性与自主搜索的学者使用。
近期研究尝试用强化学习训练大语言模型(LLM)搜索代理以解决开放域问答问题。然而,现有评估多仅关注最终答案准确率,忽视了代理如何利用外部证据进行推理和行动。我们提出SeekBench,首个通过步骤级分析响应轨迹来评估LLM搜索代理认知能力的基准。该基准包含190条专家标注的轨迹,超过1800个响应步骤,每条均附有证据标注,用于细致分析代理是否(1)在观察到的证据基础上生成推理步骤;(2)自适应地重构搜索以从低质量结果中恢复;(3)具备恰当的校准能力,正确判断当前证据是否足以给出答案。
原文摘要 · Abstract (English)
Recent work has explored training Large Language Model (LLM) search agents with reinforcement learning (RL) for open-domain question answering (QA). However, most evaluations focus solely on final answer accuracy, overlooking how these agents reason with and act on external evidence. We introduce SeekBench, the first benchmark for evaluating the \textit{epistemic competence} of LLM search agents through step-level analysis of their response traces. SeekBench comprises 190 expert-annotated traces with over 1,800 response steps generated by LLM search agents, each enriched with evidence annotations for granular analysis of whether agents (1) generate reasoning steps grounded in observed evidence, (2) adaptively reformulate searches to recover from low-quality results, and (3) have proper calibration to correctly assess whether the current evidence is sufficient for providing an answer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。