评测大模型引用网页的来源质量,发现多数引用不靠谱。
SourceBench: Can AI Answers Reference Quality Web Sources?
- 构建多维度评估框架,从内容与页面信号判断引用质量
- 测试8个大模型和3个AI搜索工具,发现平均引用可信度不足50%
- 适合关注生成式AI可靠性、信息溯源的研究者和开发者
大型语言模型(LLMs)越来越多地通过引用网络来源回答问题,但现有评估侧重答案正确性而非证据质量。我们提出SourceBench,一个涵盖100个真实世界查询的基准,覆盖信息性、事实性、论证性、社交性和购物意图。该基准采用八项指标框架,涵盖内容质量(相关性、准确性、客观性)与页面级信号(如新鲜度、权威性/可问责性、清晰度)。包含人工标注数据集,并使用校准后的基于LLM的评估器,其判断与专家高度一致。我们在3996个引用来源上评估了8个LLMs、Google Search及3个AI搜索工具,并展开进一步实验以理解结果。整体工作揭示四个关键新洞见,可指导未来生成式AI与网络搜索研究方向。
原文摘要 · Abstract (English)
Large language models (LLMs) increasingly answer queries by citing web sources, but existing evaluations emphasize answer correctness rather than evidence quality. We introduce SourceBench, a benchmark for measuring the quality of cited web sources across 100 real-world queries spanning informational, factual, argumentative, social, and shopping intents. SourceBench uses an eight-metric framework covering content quality (content relevance, factual accuracy, objectivity) and page-level signals (e.g., freshness, authority/accountability, clarity), and includes a human-labeled dataset with a calibrated LLM-based evaluator that matches expert judgments closely. We evaluate eight LLMs, Google Search, and three AI search tools over 3996 cited sources using SourceBench and conduct further experiments to understand the evaluation results. Overall, our work reveals four key new insights that can guide future research in the direction of GenAI and web search.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。