首次系统识别出大量由大模型主导的网页,揭示其规模与增长趋势。
DeGenTWeb: A First Look at LLM-dominant Websites
- 通过多页面检测聚合,精准识别大模型主导的网站
- 在Common Crawl和Bing搜索中发现大模型网页占比高且持续上升
- 最新大模型使内容识别更难,凸显检测挑战
近期不少媒体报道称大语言模型(LLMs)生成的内容正占领网络,但这些结论缺乏代表性样本和透明方法。为避免误判人类内容,我们发现现有大模型内容检测器性能远低于宣传水平。为此,我们提出DeGenTWeb,系统识别大模型主导的网站——即内容主要由大模型生成、人工干预极少的站点。我们改进了网页级大模型内容检测方法,并通过聚合多个页面结果实现站点级别的准确分类。利用DeGenTWeb,我们在Common Crawl数据和Bing搜索结果中均发现大模型主导网站高度普遍存在,且比例随时间持续增长。同时表明,面对最新大模型能力,持续准确识别此类网站仍具挑战性。
原文摘要 · Abstract (English)
Many recent news reports have claimed that content generated by large language models (LLMs) is taking over the web. However, these claims are typically not based on a representative sample of the web and the methodology underlying them is often opaque. Moreover, when aiming to minimize the chances of falsely attributing human-authored content to LLMs, we find that detectors of LLM-generated text perform much worse than advertised. Consequently, we lack an understanding of the true prevalence and characteristics of LLM content on the web. We describe DeGenTWeb which systematically identifies LLM-dominant websites: sites whose content has been generated using LLMs with little human input. We show how to adapt detectors of LLM-generated text for use on web pages, and how to aggregate detection results from multiple pages on a site for accurate site-level categorization. Using DeGenTWeb, we find that LLM-dominant sites are highly prevalent both in data from Common Crawl and in Bing's search results, and that this share is growing over time. We also show that continuing to accurately identify such sites appears challenging given the capabilities of the latest LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。