为生成式AI数据集设计合规评分框架,追踪来源保障安全透明
Compliance Rating Scheme: A Data Provenance Framework for Generative AI Datasets
- 基于数据溯源技术构建合规评分体系,评估数据透明性与安全性
- 可对现有数据集进行合规评分,也能指导新数据采集的负责任实践
- 开源工具包支持无缝集成到训练流程,适合关注数据伦理的研究者
生成式人工智能(GAI)近年来快速发展,部分得益于大规模开源数据集的可用性。然而,这些数据集常采用无限制且不透明的数据收集方式。尽管多数研究聚焦于GAI模型本身,但其数据集创建过程中的伦理与法律问题往往被忽视。随着数据集在网上传播、修改和再生产,其来源、合法性与安全性信息常丢失。为此,我们提出合规评分方案(Compliance Rating Scheme, CRS),一种用于评估数据集在透明性、问责性与安全性方面合规性的框架。同时发布一个基于数据溯源技术的开源Python库,支持该框架的实现,可无缝集成至现有数据处理与AI训练流程中。该工具兼具反应式与前瞻性功能:既可评估已有数据集的CRS分数,也能指导负责任的数据抓取与新数据集构建。
原文摘要 · Abstract (English)
Generative Artificial Intelligence (GAI) has experienced exponential growth in recent years, partly facilitated by the abundance of large-scale open-source datasets. These datasets are often built using unrestricted and opaque data collection practices. While most literature focuses on the development and applications of GAI models, the ethical and legal considerations surrounding the creation of these datasets are often neglected. In addition, as datasets are shared, edited, and further reproduced online, information about their origin, legitimacy, and safety often gets lost. To address this gap, we introduce the Compliance Rating Scheme (CRS), a framework designed to evaluate dataset compliance with critical transparency, accountability, and security principles. We also release an open-source Python library built around data provenance technology to implement this framework, allowing for seamless integration into existing dataset-processing and AI training pipelines. The library is simultaneously reactive and proactive, as in addition to evaluating the CRS of existing datasets, it equally informs responsible scraping and construction of new datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。