为企业级大模型设计了真实场景下的指令遵循评测基准。
FireBench: Evaluating Instruction Following in Enterprise and API-Driven LLM Applications
- 基于真实企业与API使用场景构建评测数据集。
- 覆盖2400+样本,测试11个模型在6大能力维度的表现。
- 适合评估大模型在客服、编程等专业任务中的可靠性。
指令遵循对部署于企业及API驱动环境的大模型至关重要,严格遵守输出格式、内容约束和流程要求是实现可靠大模型辅助工作流的基础。然而,现有指令遵循评测主要针对聊天助手的自然语言生成需求,未能反映企业用户的真实场景。为此,我们提出FireBench,一个基于真实企业与API使用模式的大模型指令遵循评测基准。该基准涵盖信息抽取、客户支持、编码代理等多样化应用,评估六大核心能力维度,包含超过2,400个样本。我们评估了11个大模型,并揭示其在企业场景下的指令遵循表现。FireBench已开源(fire-bench.com),旨在帮助用户评估模型适用性,支持开发者诊断性能瓶颈,并欢迎社区贡献。
原文摘要 · Abstract (English)
Instruction following is critical for LLMs deployed in enterprise and API-driven settings, where strict adherence to output formats, content constraints, and procedural requirements is essential for enabling reliable LLM-assisted workflows. However, existing instruction following benchmarks predominantly evaluate natural language generation constraints that reflect the needs of chat assistants rather than enterprise users. To bridge this gap, we introduce FireBench, an LLM instruction following benchmark grounded in real-world enterprise and API usage patterns. FireBench evaluates six core capability dimensions across diverse applications including information extraction, customer support, and coding agents, comprising over 2,400 samples. We evaluate 11 LLMs and present key findings on their instruction following behavior in enterprise scenarios. We open-source FireBench at fire-bench.com to help users assess model suitability, support model developers in diagnosing performance, and invite community contributions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。