arXiv:2609.05776cs.AI2026-09

构建企业数据智能评估基准,测试模型融合业务知识与计算的能力。

DI-Bench: Systematically Generating In-Domain Data Intelligence Benchmarks for Enterprise Agents

论文配图:DI-Bench: Systematically Generating In-Domain Data Intelligence Benchmarks for Enterprise Agents
图 1 · 摘自论文原文
  • 通过构建数据资产关联图生成真实任务,结合结构化数据与业务文档。
  • 生成731个任务,涵盖知识检索、分析计算和规则推理,其中计算任务准确率仅32%。
  • 适合评估企业级AI代理在复杂业务场景下的综合理解能力。

在领域特定基准上评估企业代理至关重要,但公开基准很少检验代理是否能将业务知识与分析计算相结合,而手动构建此类基准成本高昂。我们提出DI-Bench,一个用于生成数据智能(DI)真实评估基准的流水线。为模拟需要计算与知识检索的现实DI任务,DI-Bench在数据表、维度、指标和文档之间构建资产关联图,形成涉及结构化数据与相关知识的问题。真实答案通过查询执行获得,随后利用大模型生成并验证问题。应用于两个公开数据集,该流水线生成了包含731个任务的基准,覆盖知识检索、分析计算和规则驱动推理。为展示基准的区分能力和难度,我们评估了四个模型,发现当检索到的业务规则影响计算时,模型准确率仅为32%。

原文摘要 · Abstract (English)

Evaluating enterprise agents on domain-specific benchmarks is critical, yet public benchmarks rarely evaluate whether agents can integrate business knowledge with analytical computation, and constructing such benchmarks manually is costly. We present DI-Bench, a pipeline for generating realistic benchmarks for data intelligence (DI), the practice of extracting insights from large volumes of enterprise data. To emulate realistic DI tasks that require both computation and knowledge retrieval, DI-Bench builds an artifact linkage graph over data tables, dimensions, metrics, and documents to form questions involving structured data and associated knowledge. Ground truth answers are derived via query execution, followed by LLM question generation and validation. Applied to two public datasets, the pipeline produces a 731-task benchmark covering knowledge retrieval, analytical computation, and rule-grounded reasoning. To show the discriminatory capability and difficulty of the benchmark, we evaluate four models, revealing a substantial finding: models achieve only 32% accuracy when doing computational tasks where retrieved business rules modify the computation.

数据智能企业代理评估基准知识融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。