AI生成内容泛滥正导致搜索系统失效,引发信息检索崩溃。
Retrieval Collapses When AI Pollutes the Web
- AI内容主导搜索结果,导致来源单一化。
- 67%污染时检索结果超80%被污染,但准确率仍看似稳定。
- 大模型排名器比传统方法更能过滤恶意内容,适合防御性系统设计。
网络上AI生成内容的快速蔓延对信息检索构成结构性风险,因搜索引擎和检索增强生成(RAG)系统越来越多地依赖大语言模型(LLMs)产生的证据。我们将其定义为检索崩溃(Retrieval Collapse),一种两阶段过程:(1) AI生成内容主导搜索结果,削弱来源多样性;(2) 低质量或对抗性内容渗入检索管道。通过控制实验分析高质量SEO内容与对抗性内容的影响。在SEO场景中,67%的池污染导致超过80%的暴露污染,形成表面健康但同质化的状态,尽管依赖合成来源,答案准确率仍保持稳定。而在对抗性污染下,传统基线如BM25暴露约19%有害内容,而基于LLM的排序器表现出更强抑制能力。研究揭示了检索系统悄然转向合成证据的风险,亟需具备检索意识的策略,防止网络基础系统陷入质量持续下降的自我强化循环。
原文摘要 · Abstract (English)
The rapid proliferation of AI-generated content on the Web presents a structural risk to information retrieval, as search engines and Retrieval-Augmented Generation (RAG) systems increasingly consume evidence produced by the Large Language Models (LLMs). We characterize this ecosystem-level failure mode as Retrieval Collapse, a two-stage process where (1) AI-generated content dominates search results, eroding source diversity, and (2) low-quality or adversarial content infiltrates the retrieval pipeline. We analyzed this dynamic through controlled experiments involving both high-quality SEO-style content and adversarially crafted content. In the SEO scenario, a 67\% pool contamination led to over 80\% exposure contamination, creating a homogenized yet deceptively healthy state where answer accuracy remains stable despite the reliance on synthetic sources. Conversely, under adversarial contamination, baselines like BM25 exposed $\sim$19\% of harmful content, whereas LLM-based rankers demonstrated stronger suppression capabilities. These findings highlight the risk of retrieval pipelines quietly shifting toward synthetic evidence and the need for retrieval-aware strategies to prevent a self-reinforcing cycle of quality decline in Web-grounded systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。