arXiv:2608.20851cs.SEcs.AI2026-08

评测智能体在企业系统语言中的工程能力,发现通用模型表现不一。

BC-Bench: Evaluating Agentic Engineering in a Domain-Specific Language for ERP

  • 构建101个真实企业系统任务的专用评测集
  • 通用模型在特定领域表现差异大,迁移效果有限
  • 支持代码、测试生成与视觉上下文理解,贴近实际开发

智能体工程系统在通用基准上表现优异,但在企业资源规划(ERP)领域的领域特定语言(DSL)中仍缺乏充分评估。我们提出BC-Bench,一个针对Microsoft Dynamics 365 Business Central所用AL语言的真实世界任务评测基准。该基准包含从两个微软生产仓库中提取的101个手工标注任务,反映真实的ERP开发流程。基于SWE-Bench方法,我们解决了AL生态中公共资源少、环境配置复杂等独特挑战。除生成功能性代码外,BC-Bench还评估测试生成能力,并支持多模态问题陈述(常见视觉上下文)。我们在两个代理框架下评估多个前沿模型,采用多轮运行指标以应对非确定性。在缺陷修复任务中,不同模型间的解决率差异大于两个代理框架间的差异;通用基准上的改进并未一致迁移到AL场景。结果凸显了领域专用评测的重要性。

原文摘要 · Abstract (English)

Agentic engineering systems have shown strong performance on general-purpose benchmarks, yet their effectiveness in enterprise resource planning (ERP) domain-specific languages (DSLs) remains underexplored. We introduce BC-Bench, a benchmark designed to evaluate agentic engineering on real-world tasks in AL, the DSL for Microsoft Dynamics 365 Business Central. BC-Bench comprises 101 manually curated tasks extracted from two Microsoft-owned production repositories, reflecting authentic ERP development workflows. Adapting the SWE-Bench methodology, we address the unique constraints of the AL ecosystem---including limited public resources and complex environment provisioning. Beyond generating functional code, BC-Bench evaluates test generation and supports multimodal problem statements where visual context is commonly present. We evaluate multiple frontier models across two agent harnesses, utilizing multi-run metrics to account for nondeterminism. In the Bug Fixing category, under our evaluated settings, between-model differences in resolution rate are larger than differences between the two evaluated agent harnesses, and improvements reported on general-purpose benchmarks do not consistently transfer to AL. These results highlight the need for domain-specific evaluation.

智能体工程企业系统代码生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。