arXiv:2509.14436cs.IRcs.AI2025-09被引 4

研究生成式搜索如何影响网页内容风格与信息呈现。

When Content is Goliath and Algorithm is David: The Style and Semantic Effects of Generative Search Engine

  • 通过对比实验发现生成式搜索偏爱可预测性高、语义相似的内容。
  • 网站内容经LLM优化后,反而提升了AI摘要的信息多样性。
  • 高学历用户完成任务更快,低学历用户获得更丰富信息。

生成式搜索引擎(GEs)利用大语言模型(LLMs)生成带网页引用的摘要,开辟了新的流量获取渠道,深刻改变了搜索引擎优化格局。我们通过与谷歌生成式及传统搜索平台交互,收集了约一万个网站的数据。实证分析显示,生成式搜索倾向于引用对底层LLM更具可预测性且来源间语义相似度更高的内容。通过使用检索增强生成(RAG)API的受控实验,证明这种引用偏好源于LLM对与其生成表达模式一致内容的内在倾向。为探索LLM在优化网站内容中的应用,我们进一步实验发现,网站所有者使用LLM润色内容后,反而提升了AI摘要中的信息多样性。最后,为评估LLM引发的信息增长对用户的影响,我们设计生成式搜索平台,并招募Prolific参与者进行随机对照实验,任务为信息获取与写作。结果表明:高学历用户输出信息多样性无显著变化,但任务完成时间显著缩短;低学历用户主要受益于输出信息密度提升,完成时间基本不变。

原文摘要 · Abstract (English)

Generative search engines (GEs) leverage large language models (LLMs) to deliver AI-generated summaries with website citations, establishing novel traffic acquisition channels while fundamentally altering the search engine optimization landscape. To investigate the distinctive characteristics of GEs, we collect data through interactions with Google's generative and conventional search platforms, compiling a dataset of approximately ten thousand websites across both channels. Our empirical analysis reveals that GEs exhibit preferences for citing content characterized by significantly higher predictability for underlying LLMs and greater semantic similarity among selected sources. Through controlled experiments utilizing retrieval augmented generation (RAG) APIs, we demonstrate that these citation preferences emerge from intrinsic LLM tendencies to favor content aligned with their generative expression patterns. Motivated by applications of LLMs to optimize website content, we conduct additional experimentation to explore how LLM-based content polishing by website proprietors alters AI summaries, finding that such polishing paradoxically enhances information diversity within AI summaries. Finally, to assess the user-end impact of LLM-induced information increases, we design a generative search engine and recruit Prolific participants to conduct a randomized controlled experiment involving an information-seeking and writing task. We find that higher-educated users exhibit minimal changes in their final outputs' information diversity but demonstrate significantly reduced task completion time when original sites undergo polishing. Conversely, lower-educated users primarily benefit through enhanced information density in their task outputs while maintaining similar completion times across experimental groups.

生成式搜索大模型信息多样性用户行为

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。