对比生成式搜索与传统搜索,发现其检索方式和结果稳定性差异显著。
Characterizing Web Search in The Age of Generative AI

- 比较谷歌、OpenAI等五款生成式搜索系统与传统搜索
- 生成式搜索依赖外部知识且结果随时间波动明显
- 适合关注搜索系统可靠性与评估方法的研究者
大语言模型(LLM)催生了生成式搜索——一种将网络信息检索并合成单一连贯回答的新范式,与传统搜索返回独立网页列表有本质区别。本文系统比较了谷歌有机搜索与来自谷歌、OpenAI、Perplexity的五款生成式搜索系统。分析显示,各引擎在内部/外部知识依赖程度、来源多样性及输出稳定性方面存在显著差异。尽管生成式搜索在主题覆盖上可媲美传统搜索,但其检索路径和合成策略迥异。此外,生成式搜索输出随时间与执行次数变化,凸显鲁棒性挑战。现有评估体系未能涵盖检索行为、合成机制与稳定性等新维度,亟需发展更全面的评价方法。
原文摘要 · Abstract (English)
The advent of LLMs has given rise to generative search, a new search paradigm in which LLMs retrieve information from the web related to a query and synthesize it into a single, coherent response. This paradigm differs fundamentally from traditional web search, where results are returned as a ranked list of independent web pages. In this paper, we ask: Along what dimensions does generative search differ from traditional search? We conduct a systematic comparison between Google organic search and five generative search systems from three providers: Google, OpenAI, and Perplexity. Our analysis reveals substantial variation among engines in their reliance on internal v.s. external knowledge, source diversity, and stability. While generative systems often achieve topical coverage comparable to traditional search, they do so using markedly different retrieval footprints and synthesis strategies. We further show that the outputs of generative search can vary across time and executions, raising new challenges for robustness. Our findings demonstrate that generative search introduces new dimensions that are not captured by existing evaluation paradigms, motivating the development of evaluations that explicitly account for retrieval behavior, synthesis, and stability in generative search systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。