构建统一基准评估生成引擎排名操控,揭示攻防权衡与检测盲区。
GEO-Bench: Benchmarking Ranking Manipulation in Generative Engine Optimization
- 统一黑盒重写与白盒梯度攻击的评估协议
- 黑盒重写在提升排名上优于梯度攻击且更隐蔽
- 适用于安全研究者与平台风控团队参考
大语言模型日益用于对商品、文档和推荐结果进行排序,这使排名操纵成为公平性与信息完整性的重要隐患。现有生成引擎优化(GEO)研究虽提出多种操纵方法,但各自使用不同数据集和评估指标,导致其相对强度与可检测性不明确。本文提出GEO-Bench,一个统一协议下的排名操纵攻击评测基准,涵盖黑盒提示攻击(TAP、Zero-Shot)、白盒梯度攻击(STS、RAF、StealthRank)以及十种白帽C-SEO策略。所有方法在五个数据集上,针对固定开源排名器(Llama-3.1-8B-Instruct)进行评估,使用有效性(NRG、Success@α、Promote@α)与隐蔽性(关键词违规率、困惑度比)双重指标。结果表明:攻击的有效性与隐蔽性存在权衡;黑盒内容重写在排名提升上可媲美或超越梯度攻击,生成文本更流畅,且在部分领域能规避关键词与困惑度检测;访问模型无法预测攻击强度。通过标准化数据集、攻击实现与指标,GEO-Bench实现了跨攻击范式的首次直接比较,为检测方法发展提供支持。
原文摘要 · Abstract (English)
Large language models (LLMs) increasingly rank products, documents, and recommendations for user queries, which makes manipulating these rankings a growing concern for fairness and information integrity. Research on generative engine optimization (GEO) has produced many manipulation methods, but each is evaluated on its own dataset with its own metrics, so their relative strength and detectability stay unclear. We present GEO-Bench, a benchmark that evaluates GEO ranking-manipulation attacks under one protocol. It unifies black-box prompt-based attacks (TAP, Zero-Shot), white-box gradient-based attacks (STS, RAF, StealthRank), and ten white-hat C-SEO strategies. We score every method on five datasets against a fixed open-weight ranker (Llama-3.1-8B-Instruct), using metrics for both effectiveness (NRG, Success@α, Promote@α) and stealth (keyword violation rate, perplexity ratio). Our evaluation shows that effectiveness and stealth trade off across adversarial attacks, that black-box content rewriting matches or exceeds gradient-based attacks on rank promotion while producing more fluent text and can evade both keyword- and perplexity-based detection on some domains, and that the access model does not predict attack strength. By standardizing datasets, attack implementations, and metrics, GEO-Bench enables the first direct comparison across these attack paradigms and supports the development of detection methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。