生成式搜索正在吞噬网络内容,导致优质网页逐渐消失。
When Search Eats the Web: A Model of Corpus Erosion under Generative Extraction
- 将可爬取网页视为公共资源,分析其被抽取时的衰减机制
- 内容质量、数量和寿命同步下降,超过阈值后资源枯竭
- 多引擎竞争会加速崩溃,但合理控制可避免灾难
生成式搜索引擎(GSE)直接从抓取的网页内容中生成答案,无需返回源站访问即可获取价值(称为捕获性提取),这削弱了支撑内容生产的流量收益。为应对,发布者可能限制爬虫访问。本文将可爬取语料库建模为共用资源——可爬取公地,由三要素描述:体量、平均质量与生命周期。在发布者两种响应模式下,证明提取行为同时降低三者:发布者退出、更新失去资金、内容更易消亡。达到特定侵蚀阈值后,语料库彻底灭绝。短视的GSE可能跨越此阈值,而长期导向的GSE则能保持在阈值之下。扩展模型至多个竞争引擎,在公地稳态价值满足凹性条件下,对称均衡提取率随引擎数量增加而上升,并趋近于阈值。引入仅偏好直接答案的用户(最有利于提取的情况),证明社会最优提取率严格低于侵蚀阈值,且不超过单一引擎的可持续最优水平。最后讨论七种生存机制。
原文摘要 · Abstract (English)
Generative search engines (GSEs) answer user queries directly from crawled web content. The capture of value from the corpus without a visit returned to the source (we call this capture extraction) diverts the traffic that finances content production. In response, publishers may restrict crawler access to their websites. In this paper, we model the crawlable corpus as a common-pool resource: the crawlable commons. It is described by three quantities: volume, average quality, and lifetime. Under two types of responses of publishers we prove that extraction degrades all three at once: publishers opt out, renewal loses its funding, and content becomes more perishable. After a given erosion threshold, the corpus goes extinct. A myopic GSE can cross this threshold, a long-run oriented GSE stays below it. We extend our model to several competing engines and prove, under a concavity condition on the steady-state value of the commons, that the symmetric equilibrium extraction rate is nondecreasing in their number and converges to the threshold. Adding users who strictly prefer direct answers, the assumption most favorable to extraction, we prove that the socially optimal extraction rate lies strictly below the erosion threshold, and no higher than the single engine's sustainable optimum. Finally, we discuss seven survival mechanisms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。