arXiv:2608.30107cs.CLcs.AI2026-08

构建1.3万+数据集的国家级地图,揭示NLP数据分布不均问题

AtlasNLP: A Country-Aware Atlas of Dataset Representation in NLP

论文配图:AtlasNLP: A Country-Aware Atlas of Dataset Representation in NLP
图 1 · 摘自论文原文
  • 系统整理超1.3万份NLP数据集的国家级信息
  • 发现数据覆盖在国家与任务间严重不均
  • 提醒研究者关注语言覆盖≠地理代表

理解NLP数据集中各国人口的代表性,对识别数据缺口、指导数据采集、衡量进展和制定AI政策至关重要。然而,地理元数据极少提供,国家层面的代表性常被笼统的语言描述掩盖。我们提出AtlasNLP,一个涵盖超过13,000个NLP数据集记录的国家感知型图谱,按标准化任务类别追踪所代表的人口及其数据生产地。该图谱包含人工标注的AtlasNLP-Gold参考集和来自ACL的大型集合AtlasNLP-Core。利用此资源,我们发现:(1)数据集覆盖在国家与任务之间高度不均;(2)数据生产与代表性存在地理不对称性;(3)语言覆盖并不等同于地理代表性。这些发现揭示了当前数据文档实践中的盲点,并推动更明确的地理元数据用于国家意识型NLP评估。

原文摘要 · Abstract (English)

Understanding which countries are represented in NLP datasets is essential for identifying gaps, targeting data collection, measuring progress, and informing AI policy. However, geographic metadata is very rarely available, and country-level representation is often hidden behind broad language-level claims. We introduce AtlasNLP, a country-aware atlas of over 13,000 NLP dataset records across normalized NLP task categories, tracking both the populations represented and where datasets are produced. AtlasNLP includes AtlasNLP-Gold, a human-curated reference set, and AtlasNLP-Core, an ACL-derived large-scale collection. Using this resource, we show that (1) dataset coverage is highly uneven across countries and tasks; (2) dataset production and representation are geographically asymmetric; and (3) language coverage does not imply geographic representation. These findings reveal blind spots in current dataset documentation practices and motivate more explicit geographic metadata for country-aware NLP evaluation.

数据集分析地理偏见NLP评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。