评估生成式搜索中文章影响力的新基准,帮内容创作者看清自己被引用的程度。
CC-GSEO-Bench: A Content-Centric Benchmark for Measuring Source Influence in Generative Search Engines
- 构建包含1000+文章与5000+问答对的评测数据集,模拟真实检索场景
- 从曝光、可信度到因果影响,多维度量化文章在生成答案中的实际贡献
- 适合内容创作者、平台方了解高质量内容在生成搜索中的真实影响力
生成式搜索引擎(GSEs)从多个来源合成对话式回答,弱化了搜索排名与数字可见性的长期关联。这对内容创作者提出核心问题:如何可靠衡量一篇文章在不同意图和追问下对生成答案的影响?我们提出CC-GSEO-Bench,一个以内容为中心的评测基准,结合大规模数据集与创作者导向的评估框架。数据集包含超过1000篇源文章和超过5000个查询-文章配对,采用一对多结构支持文章级评估。通过公共QA数据集种子查询,结合有限合成增强,并仅保留源文章在后续检索中重新出现的查询,确保构建的真实性。在此基础上,我们从三个核心维度定义影响力:暴露度、忠实署名与因果影响;以及两个内容质量维度:可读性与结构、可信度与安全性。通过对每篇文章的查询聚类聚合查询级信号,总结其影响力强度、覆盖范围与稳定性,并实证刻画代表性内容模式下的影响力动态。
原文摘要 · Abstract (English)
Generative Search Engines (GSEs) synthesize conversational answers from multiple sources, weakening the long-standing link between search ranking and digital visibility. This shift raises a central question for content creators: How can we reliably quantify a source article's influence on a GSE's synthesized answer across diverse intents and follow-up questions? We introduce CC-GSEO-Bench, a content-centric benchmark that couples a large-scale dataset with a creator-centered evaluation framework. The dataset contains over 1,000 source articles and over 5,000 query-article pairs, organized in a one-to-many structure for article-level evaluation. We ground construction in realistic retrieval by combining seed queries from public QA datasets with limited synthesized augmentation and retaining only queries whose paired source reappears in a follow-up retrieval step. On top of this dataset, we operationalize influence along three core dimensions: Exposure, Faithful Credit, and Causal Impact, and two content-quality dimensions: Readability and Structure, and Trustworthiness and Safety. We aggregate query-level signals over each article's query cluster to summarize influence strength, coverage, and stability, and empirically characterize influence dynamics across representative content patterns.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。