通过特征级优化提升生成式引用可见性,兼顾内容质量。
Think Before Writing: Feature-Level Multi-Objective Optimization for Generative Citation Visibility

- 在网页特征空间而非文本层面进行多目标优化,提升可解释性。
- 在三个生成引擎上显著提高引用可见性,同时保持或提升内容质量。
- 方法适用于不同规模的语言模型,对文档级特征更敏感。
生成式问答系统通过选择性引用内容呈现信息,改变了可见性的决定方式,亟需超越传统搜索引擎优化的新方法。现有生成式引擎优化(GEO)主要依赖词元级文本重写,解释性差且难以平衡引用可见性与内容质量。我们提出 FeatGEO,一种基于特征级的多目标优化框架,将网页抽象为可解释的结构、内容和语言属性。不直接修改文本,而是优化这些特征,并由语言模型将配置转化为自然语言,实现高层优化与表层生成的解耦。在 GEO-Bench 上对三种生成引擎的实验表明,FeatGEO 持续提升引用可见性,同时保持或改善内容质量,显著优于词元级基线。进一步分析显示,引用行为更受文档级内容属性影响,而非孤立的词汇修改,且学习到的特征配置可在不同规模语言模型间泛化。
原文摘要 · Abstract (English)
Generative answer engines expose content through selective citation rather than ranked retrieval, fundamentally altering how visibility is determined. This shift calls for new optimization methods beyond traditional search engine optimization. Existing generative engine optimization (GEO) approaches primarily rely on token-level text rewriting, offering limited interpretability and weak control over the trade-off between citation visibility and content quality. We propose FeatGEO, a feature-level, multi-objective optimization framework that abstracts webpages into interpretable structural, content, and linguistic properties. Instead of directly editing text, FeatGEO optimizes over this feature space and uses a language model to realize feature configurations into natural language, decoupling high-level optimization from surface-level generation. Experiments on GEO-Bench across three generative engines demonstrate that FeatGEO consistently improves citation visibility while maintaining or improving content quality, substantially outperforming token-level baselines. Further analyses show that citation behavior is more strongly influenced by document-level content properties than by isolated lexical edits, and that the learned feature configurations generalize across language models of different scales.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。