arXiv:2605.26400cs.IRcs.AI2026-05

为结构化生成式搜索摘要设计评估框架,提升结果可信度。

Plans for Evaluating Structured Generative Search Summaries

  • 基于大模型生成含概要、分段标题和引用来源的摘要。
  • 提出可扩展的评估方案,涵盖内容准确性和结构完整性。
  • 适合关注搜索摘要质量与可信度的研究者或工程师。

我们提出一种评估框架,用于评价置于自然网络搜索结果之上的结构化生成式搜索摘要。此类摘要由大型语言模型生成,通常包含概要、若干带标题的章节以及摘要中引用的源文档列表。本文进一步描述了该框架的实施与评估计划,旨在系统化衡量摘要在信息组织、准确性与可追溯性方面的表现,为未来搜索摘要的优化提供量化依据。

原文摘要 · Abstract (English)

We propose a framework for evaluating structured generative search summaries that are placed atop organic web search results. A structured summary, generated by a large language model, typically consists of an overview, several sections with section titles, and a list of source documents that are cited within the summary. We then describe our plans for implementing and evaluating the framework.

搜索摘要评估框架大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。