首个跨工业领域与语言的代码生成评测基准,助力真实场景模型能力评估。
IndustryCode: A Benchmark for Industry Code Generation
- 构建覆盖125个工业问题的多领域多语言代码生成评测集
- Claude 4.5 Opus在子问题上达68.1%准确率,主问题42.5%
- 适合评估工业级代码生成模型的真实泛化能力
大型语言模型(LLMs)在代码生成与理解方面的进展已成为工业智能与决策优化的核心驱动力,广泛应用于金融、自动化和航空航天等领域。尽管近期成果展示了LLMs在通用代码生成上的巨大潜力,但现有评测基准大多局限于单一领域和编程语言,难以有效评估真实工业应用所需的泛化能力,也无法反映复杂工业场景下的编码要求。为此,我们提出IndustryCode,首个涵盖多个工业领域与编程语言的综合性评测基准。该基准包含从125个主要工业挑战中衍生出的579个子问题,配有严谨的问题描述与测试用例,覆盖金融、自动化、航空航天、遥感等领域,并支持MATLAB、Python、C++、Stata等多种语言。在评估中,表现最佳的模型Claude 4.5 Opus在子问题上达到68.1%的准确率,在主问题上为42.5%。该数据集及自动化评估代码将在论文接受后公开。
原文摘要 · Abstract (English)
Code generation and comprehension by Large Language Models (LLMs) have emerged as core drivers of industrial intelligence and decision optimization, finding widespread application in fields such as finance, automation, and aerospace. Although recent advancements have demonstrated the remarkable potential of LLMs in general code generation, existing benchmarks are mainly confined to single domains and languages. Consequently, they fail to effectively evaluate the generalization capabilities required for real-world industrial applications or to reflect the coding proficiency demanded by complex industrial scenarios. To bridge this gap, we introduce IndustryCode, the first comprehensive benchmark designed to span multiple industrial domains and programming languages. IndustryCode comprises 579 sub-problems derived from 125 primary industrial challenges, accompanied by rigorous problem descriptions and test cases. It covers a wide range of fields, including finance, automation, aerospace, and remote sensing-and incorporates diverse programming languages such as MATLAB, Python, C++, and Stata. In our evaluation, the top-performing model, Claude 4.5 Opus, achieved an overall accuracy of 68.1% on sub-problems and 42.5% main problems. The benchmark dataset and automated evaluation code will be made publicly available upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。