首份研究揭示大模型搜索对黑帽SEO的防御能力与新漏洞
Unveiling the Resilience of LLM-Enhanced Search Engines against Black-Hat SEO Manipulation
- 构建1000个真实黑帽网站基准,系统测试大模型搜索安全
- 传统黑帽攻击99.78%被拦截,检索阶段是主要防御层
- 新攻击策略可使操纵率翻倍,适合关注AI搜索安全的研究者
大型语言模型增强型搜索引擎(LLMSE)通过融合网络级搜索与AI摘要能力,革新了信息检索。尽管其效率优于传统引擎,但对成熟黑帽搜索引擎优化(SEO)攻击的安全性尚未被探索。本文首次系统研究针对LLMSE的SEO攻击,评估十款代表性产品(如ChatGPT、Gemini),并构建包含1000个真实黑帽SEO网站的SEO-Bench基准。测量显示,LLMSE可抵御超过99.78%的传统SEO攻击,检索阶段作为主要过滤层拦截了绝大多数恶意查询。进一步提出并评估七种新型LLMSEO攻击策略,发现现成的LLMSE易受攻击:重写查询注入与分段文本攻击使操纵率较基线翻倍。本工作提供了首个深入的LLMSE生态安全分析,为构建更鲁棒的AI驱动搜索系统提供实践洞见。相关问题已负责任地通报主要厂商。
原文摘要 · Abstract (English)
The emergence of Large Language Model-enhanced Search Engines (LLMSEs) has revolutionized information retrieval by integrating web-scale search capabilities with AI-powered summarization. While these systems demonstrate improved efficiency over traditional search engines, their security implications against well-established black-hat Search Engine Optimization (SEO) attacks remain unexplored. In this paper, we present the first systematic study of SEO attacks targeting LLMSEs. Specifically, we examine ten representative LLMSE products (e.g., ChatGPT, Gemini) and construct SEO-Bench, a benchmark comprising 1,000 real-world black-hat SEO websites, to evaluate both open- and closed-source LLMSEs. Our measurements show that LLMSEs mitigate over 99.78% of traditional SEO attacks, with the phase of retrieval serving as the primary filter, intercepting the vast majority of malicious queries. We further propose and evaluate seven LLMSEO attack strategies, demonstrating that off-the-shelf LLMSEs are vulnerable to LLMSEO attacks, i.e., rewritten-query stuffing and segmented texts double the manipulation rate compared to the baseline. This work offers the first in-depth security analysis of the LLMSE ecosystem, providing practical insights for building more resilient AI-driven search systems. We have responsibly reported the identified issues to major vendors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。