arXiv:2608.16824cs.LGcs.CR2026-08被引 1

发现并量化生成式搜索优化内容,揭示虚假信息传播风险。

GEO-Flag: Detecting and Measuring GEO-Optimized Web Content

论文配图:GEO-Flag: Detecting and Measuring GEO-Optimized Web Content
图 1 · 摘自论文原文
  • 构建3200条网页数据集,系统评估GEO检测方法有效性。
  • 提出干预配对训练法,将检测准确率提升至94.4%。
  • 在真实搜索结果中发现8.9%网页被优化,2026年高达16.36%。

生成式引擎优化(GEO)通过修改网页内容,提高其被生成式搜索引擎选中和引用的可能性,可能导致低权威或虚假信息获得过度曝光。与传统搜索不同,生成式搜索直接合成答案而非展示多方来源,加剧了信息可信度评估的难度。尽管存在此类风险,系统性检测GEO内容的方法仍不充分。本文提出 exttt{GEOFlagBench}基准,包含400个查询、4个领域、8类GEO优化器的3,200个网页实例,用于系统评估现有检测方法。最强基线的聚合F1为0.880,但方法级和作者条件分析显示其存在依赖作者特征的缺陷。为此,我们提出干预配对训练(IPT),监督检测器对GEO干预与非GEO AI润色的响应;在ModernBERT上,F1从0.862提升至0.944,最差组准确率从0.725升至0.883。进一步构建GEO门控代理系统,审计被检测页面的源层级与引用链接可验证性。最后,将完整流程部署于1,000个真实用户查询的Google Search与Gemini驱动检索结果,共分析10,095页,估算整体GEO流行度为8.90%,2026年修改的页面中达16.36%。研究为现实搜索生态中系统检测、审计与测量GEO奠定了基础。

原文摘要 · Abstract (English)

Generative Engine Optimization (GEO) modifies web content to increase its likelihood of being selected and cited by generative search engines. This can give strategically optimized pages visibility disproportionate to their authority or relevance and even make weak or false information appear well supported. Unlike conventional search, generative search synthesizes information into direct answers rather than presenting competing sources, which can further amplify these risks, as assessing source provenance and authority requires additional user interaction. Despite these concerns, systematic methods for detecting GEO-optimized webpages remain underexplored. We introduce \texttt{GEOFlagBench}, a benchmark of 3,200 web content instances spanning 400 queries, four domains, and eight GEO optimizer families, and use it to systematically evaluate existing GEO detection methods. Although the strongest baseline achieves an aggregate F1 of 0.880, method-level and authorship-conditioned evaluations reveal substantial weaknesses and potential reliance on authorship-related shortcuts. We therefore propose \emph{Intervention-Paired Training} (IPT), which supervises detector responses to GEO interventions and non-GEO AI polishing; on ModernBERT, IPT improves F1 from 0.862 to 0.944 and worst-group accuracy from 0.725 to 0.883. We develop a GEO-gated Agent system for auditing the Source Tier and verifiability of Citation URLs in detected GEO pages. Finally, we deploy the complete pipeline on released Google Search and Gemini-grounded retrieval results for 1,000 real-user queries. Across 10,095 available pages, we estimate an overall GEO prevalence of 8.90\%, reaching 16.36\% among pages modified in 2026. Our results establish a foundation for systematically detecting, auditing, and measuring GEO in real-world search ecosystems.

生成式搜索内容优化可信度检测AI审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。