32B参数代码模型,专攻芯片设计等工业级编程难题
InCoder-32B: Code Foundation Model for Industrial Scenarios
- 从头训练+工业代码精调,支持128K上下文推理
- 在9个工业任务上超越现有开源模型,通用任务表现也顶尖
- 适合芯片、嵌入式、编译器等高要求工程场景使用
近期代码大模型在通用编程任务上取得显著进展,但在需理解硬件语义、专用语言结构和严格资源约束的工业场景中性能大幅下降。为此,我们提出InCoder-32B(Industrial-Coder-32B),首个320亿参数的代码基础模型,统一覆盖芯片设计、GPU内核优化、嵌入式系统、编译器优化和3D建模等四大工业领域。通过高效架构,我们从零开始训练:先用通用代码预训练,再经精选工业代码精调,中期逐步将上下文扩展至128K tokens(使用合成工业推理数据),最后以执行验证进行后训练。在14个主流通用代码基准和9个工业基准(涵盖4个专业领域)上评估,结果表明InCoder-32B在通用任务上表现优异,同时在工业领域建立了强有力的开源基准。
原文摘要 · Abstract (English)
Recent code large language models have achieved remarkable progress on general programming tasks. Nevertheless, their performance degrades significantly in industrial scenarios that require reasoning about hardware semantics, specialized language constructs, and strict resource constraints. To address these challenges, we introduce InCoder-32B (Industrial-Coder-32B), the first 32B-parameter code foundation model unifying code intelligence across chip design, GPU kernel optimization, embedded systems, compiler optimization, and 3D modeling. By adopting an efficient architecture, we train InCoder-32B from scratch with general code pre-training, curated industrial code annealing, mid-training that progressively extends context from 8K to 128K tokens with synthetic industrial reasoning data, and post-training with execution-grounded verification. We conduct extensive evaluation on 14 mainstream general code benchmarks and 9 industrial benchmarks spanning 4 specialized domains. Results show InCoder-32B achieves highly competitive performance on general tasks while establishing strong open-source baselines across industrial domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。