arXiv:2605.15079cs.LGcs.DB2026-05被引 1

让本地数据集自动生成可发现、可管理的元数据,提升机器学习数据可用性。

Croissant Baker: Metadata Generation for Discoverable, Governable, and Reusable ML Datasets

论文配图:Croissant Baker: Metadata Generation for Discoverable, Governable, and Reusable ML Datasets
图 1 · 摘自论文原文
  • 通过命令行工具直接解析数据目录,自动创建符合Croissant标准的元数据。
  • 在140多个数据集上验证,最高达97-100%与人工标注一致,支持百万级数据规模。
  • 适合需合规管理的机构数据团队,尤其适用于大型本地数据仓库场景。

Croissant已成为机器学习数据集的元数据标准,采用基于JSON-LD的结构化格式,使数据发现、自动导入和可复现分析在不同平台间实现机器可读。尽管其采用率迅速上升,如NeurIPS已要求所有数据集投稿必须包含Croissant元数据,但实际应用中元数据生成通常始于将数据上传至公共平台,这对受监管且体量庞大的本地数据仓库不适用。为此,我们发布Croissant Baker——一款以本地优先、开源的命令行工具,通过模块化处理器注册表,直接从数据目录生成经验证的Croissant元数据。我们在超过140个数据集上进行了评估,涵盖规模达8.86亿行、374个Parquet文件的MIMIC-IV数据集。在与生产者提供的或标准推导的基准对比中,Croissant Baker在多个领域实现了97%-100%的一致性。

原文摘要 · Abstract (English)

Croissant has emerged as the metadata standard for machine learning datasets, providing a structured, JSON-LD-based format that makes dataset discovery, automated ingestion, and reproducible analysis machine-checkable across ML platforms. Adoption has accelerated, and NeurIPS now requires Croissant metadata in every submission to its dataset tracks. Yet in practice Croissant generation usually starts with uploading data to a public platform, a path infeasible for governed and large local repositories that hold much of the high-value data ML increasingly relies on. We release Croissant Baker, a local-first, open-source command-line tool that generates validated Croissant metadata directly from a dataset directory through a modular handler registry. We evaluate Croissant Baker on over 140 datasets, scaling to MIMIC-IV at 886 million rows and 374 Parquet files. On held-out comparisons against producer-authored or standards-derived ground truth, Croissant Baker reaches 97-100% agreement across multiple domains.

元数据数据治理本地优先ML数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。