arXiv:2412.17847cs.AIcs.CL2024-12被引 1

首次跨模态审计海量数据来源,揭示AI训练数据的西方中心化与透明度危机。

Bridging the Data Provenance Gap Across Text, Speech and Video

  • 系统追踪1990-2024年近4000个文本、语音、视频数据集的来源与授权
  • 超80%主流数据内容含非商业使用限制,但仅不足33%数据集受严格许可
  • 尽管语言和地域覆盖增多,但相对代表性自2013年未明显提升

人工智能进展主要依赖训练数据的规模与质量。然而,对文本之外的数据集属性的实证分析仍显不足。本文首次开展跨模态(文本、语音、视频)的纵向审计,覆盖1990至2024年间近4000个公开数据集,涵盖608种语言、798个来源、659个组织及67个国家。研究发现,多模态机器学习应用自2019年起主要依赖网络爬取、合成数据及社交媒体平台(如YouTube)作为训练数据,取代其他来源。其次,追溯数据衍生链显示,虽不到33%的数据集具有严格许可,但超过80%的主流文本、语音与视频数据中的原始内容携带非商业使用限制。最后,尽管公共数据集中语言与地理覆盖数量上升,但相对代表性自2013年以来未显著改善。本研究为理解数据溯源、使用限制与西方中心化趋势提供了生态系统级实证依据,强调透明度对负责任AI的重要性。为此,我们公开完整多模态审计结果,支持从业者追溯跨模态数据来源。

原文摘要 · Abstract (English)

Progress in AI is driven largely by the scale and quality of training data. Despite this, there is a deficit of empirical analysis examining the attributes of well-established datasets beyond text. In this work we conduct the largest and first-of-its-kind longitudinal audit across modalities--popular text, speech, and video datasets--from their detailed sourcing trends and use restrictions to their geographical and linguistic representation. Our manual analysis covers nearly 4000 public datasets between 1990-2024, spanning 608 languages, 798 sources, 659 organizations, and 67 countries. We find that multimodal machine learning applications have overwhelmingly turned to web-crawled, synthetic, and social media platforms, such as YouTube, for their training sets, eclipsing all other sources since 2019. Secondly, tracing the chain of dataset derivations we find that while less than 33% of datasets are restrictively licensed, over 80% of the source content in widely-used text, speech, and video datasets, carry non-commercial restrictions. Finally, counter to the rising number of languages and geographies represented in public AI training datasets, our audit demonstrates measures of relative geographical and multilingual representation have failed to significantly improve their coverage since 2013. We believe the breadth of our audit enables us to empirically examine trends in data sourcing, restrictions, and Western-centricity at an ecosystem-level, and that visibility into these questions are essential to progress in responsible AI. As a contribution to ongoing improvements in dataset transparency and responsible use, we release our entire multimodal audit, allowing practitioners to trace data provenance across text, speech, and video.

数据溯源负责任AI多模态数据透明

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。