arXiv:2607.14035cs.IRcs.DL2026-07综述被引 1

梳理生成引擎优化的可见性机制,揭示现有方法的局限与可复现性问题。

Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)

  • 构建多阶段可视性模型,区分发现、引用、吸收与经济结果
  • 实证发现内容位置和主题相关性最有效,通用策略效果差
  • 强调重复实验与人类验证,适合关注可信AI评估的研究者

生成引擎优化(GEO)旨在提升内容在生成引擎输出中的曝光度、被引用概率或影响力。本文综述2023年11月至2026年7月间发表的45项研究,含1篇早期预印本及相关的RAG与评估工作。我们指出,GEO并非单一排序任务,而是一个涉及搜索激活、爬取索引、检索、重排序、上下文分配、引用、显著性、事实吸收、保真度与用户行为的随机且部分可观测的流程。基础论文中被广泛引用的增益仅在其实验设定内成立,且依赖已有内容存在于固定上下文中;未能证明有机发现能力或持久流量影响。研究显示,主题相关性与上下文位置是最可复现的杠杆,通用启发式方法迁移性差,竞争可能抵消个体收益,以引用为导向的改写反而损害检索效果。商业审计还发现源内容重叠率低、运行间变异性大、保真度缺口持续存在。本文提出多阶段形式化模型、可见性向量、证据等级体系,以及基于重复测量、改写、对照、人工验证与多主体干扰的可复现协议。在该语料中,证据范围有限:已被检索的内容可因果影响其引用或使用,但无任一技术表现出稳定、纵向、跨平台对有机发现或下游行为的因果效应。

原文摘要 · Abstract (English)

Generative Engine Optimization (GEO) seeks to increase content's presence, likelihood of citation, or influence in answers produced by generative engines. Since the foundational GEO paper, the field has expanded rapidly, but terminology, metrics, and evidence standards remain heterogeneous. This critical survey reviews 45 studies selected under a November 2023-July 2026 publication window, including one earlier preprint published at EMNLP after the window opened, plus relevant RAG and evaluation work. We argue that GEO is not a single ranking task but a stochastic, partially observable pipeline spanning search activation, crawling and indexing, retrieval, reranking and context allocation, citation, prominence, factual absorption, fidelity, and user behavior. The foundational paper's widely cited gains are valid within its experimental setting but conditional on a source already being present in a fixed context; they establish neither organic discoverability nor durable traffic effects. Reviewed work indicates that topical relevance and context position are the most reproducible levers, generic heuristics transfer poorly, competition can erode individual gains, and citation-oriented rewrites can impair retrieval. Commercial audits further reveal low source overlap, substantial run-to-run variability, and persistent fidelity gaps. We contribute a multistage formal model, a visibility vector separating discoverability, citation, absorption, and economic outcomes, an evidence hierarchy, and a reproducible protocol based on repeated measurements, paraphrases, controls, human validation, and multi-actor interference. Within this corpus, the evidence is narrow: already-retrieved content can causally alter its citation or use, but no reviewed technique shows a stable, longitudinal, cross-platform causal effect on organic discoverability or downstream behavior.

生成引擎可见性优化可复现性评估方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。