构建超大规模企业问答基准,评估大模型在真实场景下的推理能力
CorporateBench: Large-Scale Q&A Benchmarking with Temporal Knowledge Bases

- 基于时间演化的知识库生成四家虚拟企业数据,覆盖12至10000名员工
- 单个语料库超23万文档,验证大模型在真实企业规模下的表现下降趋势
- 适用于企业级AI开发,填补大模型在复杂通信场景评测的空白
大型语言模型(LLMs)日益具备回答企业级文档集合中复杂问题的能力。然而,评估困难:企业不愿共享内部沟通数据,而现有合成数据集过于简单。本文提出CorporateBench(CB),一个经人工验证的多任务问答基准,其规模接近真实企业通信网络中的条件,评估语料库总文档数超过230,000份。CB通过四个合成企业(员工规模12至10,000人)在两个维度(信息抽取与知识库查询)上评估五种LLMs,每个语料库均源自随时间演化的知识库,确保跨数十万文档的逻辑一致性。实验显示,随着输入规模接近真实场景,模型性能持续下降。该基准为大模型在企业通信推理能力方面提供了关键评估指标,填补了评测体系中的重要空白。
原文摘要 · Abstract (English)
LLMs are increasingly able to answer complex questions about enterprise-scale document collections. But evaluation is hard: companies don't want to share internal communications, and synthetic datasets have been overly simple. We present CorporateBench (CB), a human-validated multi-task Q&A benchmark whose scale approaches the conditions LLMs encounter in corporate communication networks, with evaluation corpora surpassing 230,000 documents. CB evaluates LLMs across two dimensions (information extraction and knowledge base querying) through four synthetically generated firms ranging from 12 to 10,000 employees. Each corpus is sampled from a temporally evolving knowledge base describing a consistent world, guaranteeing cross-document logical consistency even across hundreds of thousands of documents. We evaluate five LLMs on CB, revealing increasingly poor performance as input size approaches realistic scales. CB provides LLM developers a metric for corporate communication reasoning, filling a crucial gap in the benchmarking ecosystem.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。