arXiv:2609.08094cs.AI2026-09

首个诊断政务搜索代理失效的框架,揭示模型主要因检索失败而非知识不足出错。

CIVI: A Framework for Diagnosing Search Agent Failures in Civic Information

  • 通过源注入消融法拆解搜索失败模式,定位根本原因。
  • 10个前沿搜索代理均不及人工基准,72.1%错误源于检索问题。
  • 涵盖多国多层级政府场景,评估引用权威来源的准确性。

大型语言模型在公共部门的应用日益广泛,但错误引导可能造成不可逆伤害。我们提出CIVI,首个用于诊断政务信息搜索代理失效的框架。其基准涵盖跨国、跨行政层级(联邦、州、地方)政府背景及联合国采纳的国际标准功能类别。评估了十个前沿搜索代理,发现无一能超越专注人类基线。除准确率外,CIVI还衡量搜索调用率、选择性无搜索准确率及引用权威政府来源的频率。为实现诊断,我们引入ARISE,将代理搜索失败分解为四种互斥模式,通过源注入消融法识别。结果表明,72.1%的失败归因于检索限制,而非模型参数知识的缺失。

原文摘要 · Abstract (English)

Large Language Models are increasingly deployed in public-sector settings, where incorrect guidance can cause irreversible harm. We introduce CIVI, the first framework for diagnosing search agent failures in civic information. Its benchmark instantiation jointly spans cross-national, interjurisdictional government contexts (federal, state, and local) and functional categories from an internationally adopted United Nations standard. We evaluate ten frontier search agents and find that none matches an attentive human baseline. Alongside accuracy, CIVI measures search invocation rate, selective no-search accuracy, and how often agents cite authoritative government sources. To perform this diagnosis, we introduce ARISE, which decomposes agentic search failures into four mutually exclusive modes, isolated via source-injection ablation. ARISE attributes 72.1% of all observed failures to retrieval-bound causes rather than to gaps in the models' parametric knowledge.

搜索代理政务信息故障诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。