对比了政府AI服务中精选资料与网络搜索的优劣,发现前者可信但覆盖少,后者覆盖广但来源不可靠。
Curated retrieval versus open web search in public AI information services: a coverage-trust trade-off

- 用精选知识库和开放网络搜索两种方式生成答案,评估其来源可信度
- 35%的网络搜索结果含不可信或无关引用,精选库仅因过时被质疑
- 提示词优化效果有限,专业领域需更透明的源头控制机制
公共机构越来越多地使用大语言模型回答公民问题,常结合精选知识库与实时网络搜索。本文对冰岛大学运营的独立政府资助服务Evrópuvefur在2026年8月29日公投前开展预发布专家评估,该服务解答关于欧盟的问题。五位领域专家对449条AI生成答案进行551次评估,依据七项标准评分并标记引用来源。比较了两种检索路径:本地精选语料库(RAG)与开放网络搜索。在187条网络搜索答案中,有35%(65条)至少有一处引用被标记为不可信或不相关;而精选来源仅因过时被质疑。网络搜索覆盖更广,但牺牲了引用质量;精选语料库虽可信,但覆盖有限,模型在无法回答时选择沉默。此外,系统从未引用冰岛最广泛使用的新闻源RÚV。提示词消融实验显示,即使在系统提示中加入可信域名列表,引用率也仅从12%提升至21%。流畅性与主题契合度无法预测引用可信度。我们主张,引用可信度是公共AI服务中可测量却常被忽视的信息质量维度,并讨论了透明性应对措施及其权衡。
原文摘要 · Abstract (English)
Public institutions increasingly use large language models (LLMs) to answer citizens' questions, often pairing a curated knowledge base with live web search, yet whether the sources behind these answers can be trusted has received little empirical scrutiny. We report a pre-launch expert evaluation of Evrópuvefur, an independent, government-funded service run by the University of Iceland that answers questions about the European Union, conducted as Iceland prepared for its referendum of 29 August 2026 on whether to resume EU accession talks. Five domain experts produced 551 evaluations of 449 AI-generated answers, scoring each against a seven-criterion quality rubric and, separately, flagging individual cited sources. We compared two retrieval paths: a curated local corpus (RAG) and open web search. In more than a third of the reviewed web-search answers (35%, 65 of 187), at least one cited source was flagged, almost always as untrustworthy or irrelevant; curated sources were flagged far less often and only for being out of date. Web search answered more questions, but at the cost of source quality; the curated corpus was trustworthy yet limited in coverage, and the model declined to respond when it fell short. The citation mix also passed over strong sources: across all 287 web-search answers, the system never cited RÚV, the public broadcaster and the country's most widely used news source. A companion prompt ablation shows how weak prompt-level steering is: a trusted-domain list in the system prompt raised the share of citations to listed domains only from 12% to 21%. Fluency and topical fit did not predict source trustworthiness. We argue that source trustworthiness is a measurable yet largely invisible dimension of information quality in public AI services, and we discuss transparency-oriented responses and their trade-offs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。