arXiv:2605.18337cs.CL2026-05被引 1

构建135亿新闻文章的高效检索系统,支持跨国家、长周期媒体研究。

Infini-News: Efficiently Queryable Access to 1.3 Billion Processed Common Crawl News Articles

  • 构建包含135亿篇文章的清洗与标注数据集,支持多语言与地理定位。
  • 通过后缀数组索引实现全文任意关键词秒级搜索。
  • 适合做社会计算、跨文化媒体分析的研究者使用。

大规模新闻语料库支撑计算社会科学和自然语言处理的广泛研究,但访问受限:商业数据库成本高且授权严格,而开源的Common Crawl CC-News需数TB存储与高强度处理。本文提出Infini-News,一个覆盖2016年8月至今全部CC-News档案的检索工具与索引系统。贡献有三:第一,提取并清洗超过135亿篇文章的文本,解析结构化元数据;第二,利用三种前沿语言识别模型(GlotLID、lingua、CommonLingua)进行语言检测,并通过多源地理归因,为222个国家中的83.4%文章确定来源国;第三,构建Infini-gram索引——基于后缀数组的结构,使研究人员可在亚秒级时间内对全库进行任意文本模式搜索。该资源显著降低长期、跨国媒体研究的门槛。

原文摘要 · Abstract (English)

Large-scale news corpora support a wide range of research in Computational Social Science and NLP, yet access remains constrained: commercial archives impose prohibitive costs and licensing restrictions, while open alternatives like Common Crawl's CC-News require terabyte-scale storage and computationally intensive processing. We present Infini-News, a retrieval toolkit and index for the entire CC-News archive from August 2016 to the latest available snapshot. Our contributions are threefold. First, we extract, clean the text, and parse the structured metadata of over 1.35B articles. Second, we enrich the corpus with language detection using three frontier language classifiers (GlotLID, lingua, and CommonLingua), and with multi-source geographic attribution that resolves a country of origin for 83.4% of articles across 222 countries. Third, we construct Infini-gram indexes: suffix-array structures that let researchers search the full archive for arbitrary text patterns in sub-second time. Together, these resources lower the barrier to longitudinal, cross-national media research.

新闻数据检索系统大语言模型社会计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。