对比维基与AI生成百科的搜索推荐,发现两者都常推意外内容。
Unexpected Knowledge: Auditing Wikipedia and Grokipedia Search Recommendations
- 用近万组词汇查询,收集七万条搜索结果进行对比分析。
- 两平台均高频推荐与查询无关的内容,且相同查询结果差异大。
- 适合关注AI内容推荐偏差、信息探索路径的研究者阅读。
百科类知识平台是用户在线探索信息的重要入口。近期推出的完全由AI生成的Grokipedia为传统权威平台如Wikipedia提供了新选择。在此背景下,搜索引擎机制在引导用户探索路径中扮演关键角色,但其在不同百科系统中的表现仍缺乏研究。本文首次对Wikipedia与Grokipedia的搜索引擎进行对比分析。我们使用近10,000个中性英文词及其子串作为查询,收集超过70,000条搜索结果,考察其语义一致性、内容重叠度及主题结构。结果发现,两个平台均频繁生成与原始查询关联较弱的结果,许多看似无害的查询也会引出意外内容。尽管存在共性,同一查询在两平台产生的推荐列表往往显著不同。通过主题标注与探索轨迹分析,我们进一步识别出内容类别呈现方式及搜索结果演化路径的系统性差异。总体而言,意外推荐是两大平台的普遍特征,尽管其主题分布和查询建议存在差异。
原文摘要 · Abstract (English)
Encyclopedic knowledge platforms are key gateways through which users explore information online. The recent release of Grokipedia, a fully AI-generated encyclopedia, introduces a new alternative to traditional, well-established platforms like Wikipedia. In this context, search engine mechanisms play an important role in guiding users exploratory paths, yet their behavior across different encyclopedic systems remains underexplored. In this work, we address this gap by providing the first comparative analysis of search engine in Wikipedia and Grokipedia. Using nearly 10,000 neutral English words and their substrings as queries, we collect over 70,000 search engine results and examine their semantic alignment, overlap, and topical structure. We find that both platforms frequently generate results that are weakly related to the original query and, in many cases, surface unexpected content starting from innocuous queries. Despite these shared properties, the two systems often produce substantially different recommendation sets for the same query. Through topical annotation and trajectory analysis, we further identify systematic differences in how content categories are surfaced and how search engine results evolve over multiple stages of exploration. Overall, our findings show that unexpected search engine outcomes are a common feature of both the platforms, even though they exhibit discrepancies in terms of topical distribution and query suggestions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。