模型风险评估不能只看模型大小,数据才是关键。
Data-Centric AI Governance: Addressing the Limitations of Model-Focused Policies
- 从数据规模和内容出发评估模型风险,而非仅关注模型类型
- 小模型用好数据也能达到大模型效果,现有监管易漏判
- 反对仓促监管,主张用量化方法建立清晰的治理框架
当前对强大AI能力的监管过度聚焦于‘基础’或‘前沿’模型,但这些术语模糊且定义不一,导致治理基础不稳。更关键的是,政策讨论常忽视模型所用数据,尽管数据与模型性能密切相关。即使(相对)小型的模型,若使用足够特定的数据集,也能实现与大型模型相当的效果。本文强调,数据规模与内容是评估模型当前及未来风险的核心要素。同时指出,过度反应式监管存在风险,并提出一种基于量化评估的能力分析路径,有望构建更清晰、简化的监管环境。
原文摘要 · Abstract (English)
Current regulations on powerful AI capabilities are narrowly focused on "foundation" or "frontier" models. However, these terms are vague and inconsistently defined, leading to an unstable foundation for governance efforts. Critically, policy debates often fail to consider the data used with these models, despite the clear link between data and model performance. Even (relatively) "small" models that fall outside the typical definitions of foundation and frontier models can achieve equivalent outcomes when exposed to sufficiently specific datasets. In this work, we illustrate the importance of considering dataset size and content as essential factors in assessing the risks posed by models both today and in the future. More broadly, we emphasize the risk posed by over-regulating reactively and provide a path towards careful, quantitative evaluation of capabilities that can lead to a simplified regulatory environment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。