arXiv:2605.12153cs.SEcs.AI2026-05

构建了包含8.3亿行代码的工业级代码数据集,支持软件工程研究。

CIDR: A Large-Scale Industrial Source Code Dataset for Software Engineering Research

  • 从工业伙伴收集4225个真实项目,覆盖75种语言
  • 含完整版本历史与持续集成等工程属性,共8.3亿行原始代码
  • 适合代码智能、质量分析等方向研究者使用

我们提出面向工业界的开发者仓库数据集CIDR,涵盖4,225个真实软件仓库,覆盖75种编程语言,总计8.32亿行原始代码(5.81亿行逻辑代码),并包含仓库级结构化元数据、完整的版本控制历史以及持续集成使用情况、自动化测试存在性等工程实践属性。所有仓库通过多阶段专门设计的采集、筛选与匿名化流程处理。此外,我们报告了一项探索性微调研究,将一个30亿参数的代码语言模型适配至CIDR,并量化其在保留企业代码上的表现。CIDR旨在支持代码智能、软件质量分析、开发工具链等软件工程研究任务。数据集访问受限制许可管控,具体申请条件见https://fermatix.ai/#Contact。

原文摘要 · Abstract (English)

We present the Curated Industrial Developer Repository (CIDR), a large-scale dataset of real-world software repositories collected from industrial partners. The dataset comprises 4,225 repositories spanning 75 programming languages, totaling 832 million raw lines of code (581 million logical lines), along with structured metadata at the repository level, full version control history, and engineering-practice attributes such as continuous integration usage and the presence of automated tests. All repositories were collected, filtered, and anonymized through a multi-stage pipeline developed specifically for this purpose. We additionally report an exploratory fine-tuning study that adapts a 3-billion-parameter code language model to CIDR and quantifies the effect on held-out enterprise code. CIDR is intended to support research in code intelligence, software quality analysis, developer tooling, and related software engineering tasks. Access to CIDR is provided under a restricted license; details on eligibility and terms are available at https://fermatix.ai/#Contact.

代码数据集软件工程工业数据代码智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。