用可量化的评分表提升数据集质量,让评估更透明高效。
Datasheets Aren't Enough: DataRubrics for Automated Quality Metrics and Accountability
- 设计数据评分框架DataRubrics,用标准指标评估数据质量
- 结合大模型评判,实现可复现、低成本的数据评估方法
- 适合数据集作者和审稿人使用,推动高质量数据研究
高质量数据集是训练和评估机器学习模型的基础,但其构建——尤其是准确的人工标注——仍面临重大挑战。许多数据集论文缺乏原创性、多样性或严格的质量控制,且这些缺陷常在同行评审中被忽略。提交材料也常缺少数据构建过程和属性的关键信息。尽管已有工具如datasheets旨在提升透明度,但它们多为描述性,缺乏标准化的可测量评估方法。会议元数据要求虽强调问责,却执行不一。为此,本文倡导将系统化、基于评分表的评估指标融入数据集评审流程,尤其在投稿量持续增长的背景下。我们探索了可扩展、低成本的合成数据生成方法,包括专用工具和大模型作为裁判(LLM-as-a-judge)策略,以支持更高效的评估。作为行动呼吁,我们提出DataRubrics——一个用于评估人类与模型生成数据集质量的结构化框架。借助大模型评估最新进展,DataRubrics提供可复现、可扩展、可操作的解决方案,使作者与审稿人能共同提升数据驱动研究的标准。相关代码已开源:https://github.com/datarubrics/datarubrics。
原文摘要 · Abstract (English)
High-quality datasets are fundamental to training and evaluating machine learning models, yet their creation-especially with accurate human annotations-remains a significant challenge. Many dataset paper submissions lack originality, diversity, or rigorous quality control, and these shortcomings are often overlooked during peer review. Submissions also frequently omit essential details about dataset construction and properties. While existing tools such as datasheets aim to promote transparency, they are largely descriptive and do not provide standardized, measurable methods for evaluating data quality. Similarly, metadata requirements at conferences promote accountability but are inconsistently enforced. To address these limitations, this position paper advocates for the integration of systematic, rubric-based evaluation metrics into the dataset review process-particularly as submission volumes continue to grow. We also explore scalable, cost-effective methods for synthetic data generation, including dedicated tools and LLM-as-a-judge approaches, to support more efficient evaluation. As a call to action, we introduce DataRubrics, a structured framework for assessing the quality of both human- and model-generated datasets. Leveraging recent advances in LLM-based evaluation, DataRubrics offers a reproducible, scalable, and actionable solution for dataset quality assessment, enabling both authors and reviewers to uphold higher standards in data-centric research. We also release code to support reproducibility of LLM-based evaluations at https://github.com/datarubrics/datarubrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。