arXiv:2609.02316cs.IRcs.CL2026-09

构建基准评估对抗生成引擎优化的防御方法,发现现有方案效果有限。

Counter-GEO-Bench: Evaluating Defenses Against Information-Distorting Generative Engine Optimization

论文配图:Counter-GEO-Bench: Evaluating Defenses Against Information-Distorting Generative Engine Optimization
图 1 · 摘自论文原文
  • 设计247组真实查询与信息保真/扭曲的GEO重写配对,形成可控评测环境。
  • 三种现成防御使攻击成功率最多下降5.7%,且均不显著;现有安全机制易被绕过。
  • 提出轻量级基线防御C-GEO Guard,降低47.6%攻击成功率,几乎无性能损失。

生成引擎优化(GEO)使内容创作者提升网页在生成式搜索引擎中的可见性,但攻击者可发布外观正常的GEO优化文档,诱导大语言模型(LLMs)检索并合成误导性回答。现有基准无法在受控条件下评估防御措施。为此,本文提出Counter-GEO-Bench,将247个经人工验证、质量筛选的查询与信息保真及信息扭曲的GEO重写配对,评估防御在三个受害LLM上的攻击成功率(ASR)、误报率和答案质量。实验表明,三种现成防御(Granite Guardian、Llama Guard 3、NeMo Self-Check Fact-Checking)最多仅降低5.7%的相对攻击成功率,其中Granite Guardian的效果不显著。安全分类护栏针对违规行为,但无法拦截流畅的信息类误导内容。为此,提出轻量级基准防御C-GEO Guard,实现47.6%的相对攻击成功率降低,同时几乎无实用损失,证明该威胁可防御。

原文摘要 · Abstract (English)

Generative engine optimization (GEO) enables content producers to increase the visibility of their web pages in generative search engines, but the same techniques can deliver targeted misinformation when adversaries publish ordinary-looking GEO-optimized documents that victim large language models (LLMs) retrieve and synthesize into distorted answers. No existing benchmark evaluates defenses against this threat under controlled conditions. Therefore, we present Counter-GEO-Bench, a defense benchmark that pairs 247 human-verified, quality-gated queries with information-preserving and information-distorting GEO rewrites, and evaluates defenses on attack success rate (ASR), false positive rate, and answer quality across three victim LLMs. Under Counter-GEO-Bench, three off-the-shelf defenses (Granite Guardian, Llama Guard 3, and NeMo Self-Check Fact-Checking) reduce ASR by at most 5.7% relative, while Granite Guardian's reduction is not statistically significant. Safety-taxonomy guardrails target policy violations, while GEO misinformation passes through them as fluent informational content. To this end, a lightweight benchmark baseline, C-GEO Guard, is proposed, reducing ASR by 47.6% relative with near-zero utility loss, which proves threat tractable.

生成对抗LLM安全信息污染防御评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。