arXiv:2412.03824cs.AI2024-12被引 5

将数据视为治理前沿AI的工具,提出五项新政策机制。

Towards Data Governance of Frontier AI Models

  • 从数据源头入手,设计可检测滥用的探针令牌和自动化过滤机制。
  • 要求模型开发者和供应商强制报告训练数据,提升透明度。
  • 适合关注AI风险管控的政策制定者与技术监管机构参考。

数据是训练和微调当前前沿人工智能模型及未来发展模型的关键。现有学术、法律与监管工作主要聚焦数据对消费者和创作者的直接危害,如隐私泄露、版权侵权、偏见与歧视。本文则转向较少被关注的问题:数据如何成为前沿AI模型的新治理能力来源。这一‘前沿数据治理’视角为监控和缓解高级AI模型的风险(尤其是其规模化后获得特定危险能力)开辟了新路径。然而,前沿数据治理面临根本挑战:数据具有非竞争性、通常不可排除、易复制且日益可合成的特性。尽管如此,本文提出针对数据供应链中关键角色——数据生产者、聚合者、模型开发者和数据供应商——的一系列政策机制。简要概述15种治理手段,重点介绍其中五项未被充分探索的政策建议,包括:生产者开发探测令牌以识别未经授权使用;对预训练与后训练数据集进行(自动)内容过滤以清除恶意信息;要求开发者与供应商履行强制性数据集报告义务;加强数据集与数据生成算法的安全性;以及对供应商实施‘了解你的客户’(KYC)要求。通过将数据不仅视为潜在危害源,更视为关键治理杠杆,本研究旨在为前沿AI模型的治理与监管提供新的政策工具。

原文摘要 · Abstract (English)

Data is essential to train and fine-tune today's frontier artificial intelligence (AI) models and to develop future ones. To date, academic, legal, and regulatory work has primarily addressed how data can directly harm consumers and creators, such as through privacy breaches, copyright infringements, and bias and discrimination. Our work, instead, focuses on the comparatively neglected question of how data can enable new governance capacities for frontier AI models. This approach for "frontier data governance" opens up new avenues for monitoring and mitigating risks from advanced AI models, particularly as they scale and acquire specific dangerous capabilities. Still, frontier data governance faces challenges that stem from the fundamental properties of data itself: data is non-rival, often non-excludable, easily replicable, and increasingly synthesizable. Despite these inherent difficulties, we propose a set of policy mechanisms targeting key actors along the data supply chain, including data producers, aggregators, model developers, and data vendors. We provide a brief overview of 15 governance mechanisms, of which we centrally introduce five, underexplored policy recommendations. These include developing canary tokens to detect unauthorized use for producers; (automated) data filtering to remove malicious content for pre-training and post-training datasets; mandatory dataset reporting requirements for developers and vendors; improved security for datasets and data generation algorithms; and know-your-customer requirements for vendors. By considering data not just as a source of potential harm, but as a critical governance lever, this work aims to equip policymakers with a new tool for the governance and regulation of frontier AI models.

AI治理数据安全政策机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。