arXiv:2608.02949cs.AIcs.DB2026-08

为拉美AI建设数据基础设施,解决数据分散与供应不足问题

On the missing data layer and a potential solution

  • 按任务-领域-语言结构组织数据,构建可发现的共享体系
  • 现有拉美数据总量远低于前沿AI需求,需系统性整合
  • 适合关注区域AI基建、数据治理与多语言模型研究者

拉丁美洲缺失AI发展的两大基础层:数据集层与基准测试层。本文聚焦数据集层,指出其面临发现难与供给不足的双重挑战。尽管拉美已有若干AI数据集,但分布于多个平台且无统一索引;即便实现完美索引,总量仍远低于前沿AI研发所需。为此,我们提出DataHub:一种以任务为导向的数据基础设施,采用/<task>/<domain>/<language>的本体组织方式,集成数据发现、元数据管理、贡献机制、许可协议与重用支持等功能。

原文摘要 · Abstract (English)

Latin America is missing two foundational layers of AI infrastructure: the dataset layer and the benchmark layer. This paper targets the dataset layer. The dataset layer faces two compounding problems: discovery and supply. Latin American AI datasets exist but are scattered across platforms with no shared index. Even with perfect indexing, the total volume would remain far below what frontier AI development requires. We propose DataHub: a task-first data infrastructure organized through the ontology /<task?>/<domain?>/<language?>, with mechanisms for dataset discovery, metadata, contribution, licensing, and reuse.

数据基础设施拉美AI多语言模型数据治理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。