分析AI问答引擎引用网页的规律,给出提升被引用率的实用策略。
AI Answer Engine Citation Behavior An Empirical Analysis of the GEO16 Framework
- 构建16个维度的GEO评分框架,量化网页质量并预测引用概率。
- 发现网页质量得分高于0.7且满足至少12个维度时,引用率显著提升。
- 适合内容创作者和SaaS企业优化网页以提高被AI引擎采纳机会。
AI问答引擎通过生成回答并引用网页来源,正日益成为获取领域知识的主要途径。本文提出GEO-16框架,将页面质量信号转化为16个支柱评分,并生成0到1之间的标准化GEO分数。基于70个产品意图提示,收集了三个引擎(Brave Summary、Google AI Overviews、Perplexity)共1,702条引用,审计了1,100个唯一网址。结果显示,引擎在引用网页的质量上存在差异,其中元数据、新鲜度、语义化HTML和结构化数据等支柱与引用高度相关。使用带领域聚类标准误的逻辑回归模型表明,整体页面质量是引用的重要预测因子;简单阈值(如G ≥ 0.70且至少12个支柱达标)可显著提升引用率。研究还报告了各引擎对比、垂直领域效应、阈值分析与诊断结果,并转化为出版者的实用指南。研究为观察性研究,聚焦英文B2B SaaS页面,讨论了局限性、效度威胁与可复现性问题。
原文摘要 · Abstract (English)
AI answer engines increasingly mediate access to domain knowledge by generating responses and citing web sources. We introduce GEO-16, a 16 pillar auditing framework that converts on page quality signals into banded pillar scores and a normalized GEO score G that ranges from 0 to 1. Using 70 product intent prompts, we collected 1,702 citations across three engines (Brave Summary, Google AI Overviews, and Perplexity) and audited 1,100 unique URLs. In our corpus, the engines differed in the GEO quality of the pages they cited, and pillars related to Metadata and Freshness, Semantic HTML, and Structured Data showed the strongest associations with citation. Logistic models with domain clustered standard errors indicate that overall page quality is a strong predictor of citation, and simple operating points (for example, G at least 0.70 combined with at least 12 pillar hits) align with substantially higher citation rates in our data. We report per engine contrasts, vertical effects, threshold analysis, and diagnostics, then translate findings into a practical playbook for publishers. The study is observational and focuses on English language B2B SaaS pages; we discuss limitations, threats to validity, and reproducibility considerations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。