诊断并修复生成引擎中引用失败问题,提升内容可见性。
Diagnosing and Repairing Citation Failures in Generative Engine Optimization
- 构建引用失败模式分类体系,定位文档未被引用的原因。
- 提出AgentGEO系统,仅修改5%内容即实现40%以上引用率提升。
- 发现通用优化可能损害长尾内容,适合关注公平信息访问的研究者。
生成引擎优化(GEO)旨在提升AI生成内容中文档的可见性。然而,现有方法仅衡量文档对响应的贡献度,而非实际驱动流量回流的引用机制。同时,这些方法采用统一重写规则,无法诊断个别文档未被引用的原因。本文提出一种诊断式GEO方法,通过分析为何文档未能被引用并针对性干预。我们构建了统一框架:(1) 首个覆盖引用流程各阶段的引用失败模式分类体系;(2) AgentGEO——一个智能体系统,利用该分类诊断故障、从工具库中选择精准修复策略,并迭代直至成功引用;(3) 以文档为中心的基准测试,评估优化效果在未见查询上的泛化能力。AgentGEO相比基线实现超40%相对引用率提升,且仅修改5%内容;而基线仅达25%。分析显示,通用优化可能伤害长尾内容,部分文档的引用障碍无法仅靠优化解决,这对人工智能中介信息获取中的公平性具有重要启示。
原文摘要 · Abstract (English)
Generative Engine Optimization (GEO) aims to improve content visibility in AI-generated responses. However, existing methods measure contribution-how much a document influences a response-rather than citation, the mechanism that actually drives traffic back to creators. Also, these methods apply generic rewriting rules uniformly, failing to diagnose why individual document are not cited. This paper introduces a diagnostic approach to GEO that asks why a document fails to be cited and intervenes accordingly. We develop a unified framework comprising: (1) the first taxonomy of citation failure modes spanning different stages of a citation pipeline; (2) AgentGEO, an agentic system that diagnoses failures using this taxonomy, selects targeted repairs from a corresponding tool library, and iterates until citation is achieved; and (3) a document-centric benchmark evaluating whether optimizations generalize across held-out queries. AgentGEO achieves over 40% relative improvement in citation rates while modifying only 5% of content, compared to 25% for baselines. Our analysis reveals that generic optimization can harm long-tail content and some documents face challenges that optimization alone cannot fully address-findings with implications for equitable visibility in AI-mediated information access.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。