arXiv:2609.06811cs.IRcs.CL2026-09

GEO可见性评分受提示语设计影响,需明确评估场景与权重规则。

Measuring GEO Visibility: Prompt Corpora Define the Answer Market

  • 通过提示语和权重构建答案市场,定义GEO评分范围
  • 同一答案因指令不同可得不同评分,凸显评分主观性
  • 适用于评估模型输出可信度的学者或平台设计者

GEO(生成式引擎优化)可见性分数聚合生成回答中源内容的出现、引用或品牌提及。提示语语料库决定评估情境,权重则决定其相对重要性。二者共同构成无需反映真实用户需求的“答案市场”。提示语措辞可改变检索结果、竞争源及生成内容。评分需识别特定出现、引用或提及。若由语言模型执行此任务,其指令可改变对不变回答的评分。本文批判性综述这些选择如何定义GEO分数所衡量的内容,参考指标有效性、总调查误差与信息检索评估研究。框架明确情境标注、提示格式、执行条件、权重与评分规则。当权重未知或待定,框架报告可接受得分集合。区分与目标人群数据和假设一致的得分(部分识别)与权重设定差异导致的变异(规范敏感性)。单一引用不等于贡献确立。文章定义在受控文档环境中对比有无源生成回答的比较方法,区别于全引擎干预中的竞争源实验。框架支持可复现计算,未报告新实验;其普遍实证有效性仍待验证。

原文摘要 · Abstract (English)

GEO (generative engine optimization) visibility scores aggregate source appearances, citations, or brand mentions in generated answers. The prompt corpus selects the situations evaluated, while weights determine their relative importance. Together they define an "answer market" that need not represent actual user demand. Prompt wording can alter retrieval, competing sources, and generated answers. Scoring then requires identifying the appearances, citations, or mentions of interest. If a language model performs this task, its instruction can change the score assigned to an unchanged answer. Our critical survey examines how these choices help define what a GEO score measures. It draws on research into whether indicators measure the intended phenomenon, total survey error, and information retrieval evaluation. The framework specifies situation annotation, prompt formulations, execution conditions, weights, and scoring rules. When weights are unknown or remain to be chosen, the framework reports sets of admissible scores. It distinguishes values compatible with data and assumptions about a target population (partial identification) from variation across weighting conventions (normative sensitivity). A citation alone does not establish a source's contribution. The article defines a comparison of answers generated with and without a source in a controlled documentary context, distinct from an intervention on the full engine with competing sources. The framework is supported by reproducible calculations. No new experiments are reported; its general empirical validity remains to be assessed.

GEO评分提示工程模型评估可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。