用可验证证书替代传统代码数据集卡片,提升质量可信度
SIEVE: Towards Verifiable Certification for Code-datasets
- 将每项数据属性检查转为机器可读的可信证书
- 支持随时验证的统计置信边界,确保结果可审计
- 适合依赖高质量代码数据的研究团队与工业界
代码智能体和实证软件工程依赖公开代码数据集,但这些数据集缺乏可验证的质量保障。静态的‘数据集卡片’虽提供信息,却不可审计,也无法提供统计保证,难以确证数据质量。各团队各自搭建临时清洗流程,导致工作分散、成本上升。本文提出 SIEVE,一个社区驱动的框架,将每项属性检查转化为可机器读取、可验证的可信证书,并具备任意时间有效的统计置信界。我们规划了研究路线图,旨在推动 SIEVE 成熟化,以可随时验证的认证取代叙述性卡片。这一转变有望降低质量保障成本,增强对代码数据集的信任。
原文摘要 · Abstract (English)
Code agents and empirical software engineering rely on public code datasets, yet these datasets lack verifiable quality guarantees. Static 'dataset cards' inform, but they are neither auditable nor do they offer statistical guarantees, making it difficult to attest to dataset quality. Teams build isolated, ad-hoc cleaning pipelines. This fragments effort and raises cost. We present SIEVE, a community-driven framework. It turns per-property checks into Confidence Cards-machine-readable, verifiable certificates with anytime-valid statistical bounds. We outline a research plan to bring SIEVE to maturity, replacing narrative cards with anytime-verifiable certification. This shift is expected to lower quality-assurance costs and increase trust in code-datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。