构建企业级槽位填充基准,揭示大模型真实场景下的严重短板。
ESF-Bench: Benchmarking Challenging Slot-Filling Scenarios for Real-World Enterprise Applications

- 基于8个真实业务领域,设计810组多轮对话样本。
- 顶级模型在该任务上仅20.7%的样本成功提取槽位。
- 适合研究企业AI落地难题与模型鲁棒性提升者。
大型语言模型(LLMs)的快速发展推动了其在企业中的广泛应用。然而,在真实场景部署中,复杂的系统约束和不可预测的用户行为带来了独特挑战。其中,槽位填充对将非结构化输入转化为可操作的结构化数据至关重要。本文提出ESF-Bench,一个包含810个多轮样本和6530个槽位的挑战性企业级槽位填充基准,覆盖8个独特领域。该基准基于在真实企业部署中观察到的57种最复杂槽位填充场景的分类体系构建,揭示了当前最先进LLMs的显著局限:GPT-OSS-120b仅能成功提取20.7%的样本槽位。为支持后续研究,我们已将数据集、分类体系及评估代码公开于GitHub。
原文摘要 · Abstract (English)
The rapid rise of large language models (LLMs) has driven transformative adoption across enterprises. However, deploying these models in real-world settings presents unique challenges due to complex system constraints and unexpected user behaviors. Among these applications, slot filling is essential for converting unstructured input into structured, actionable data. In this work, we introduce ESF-Bench, a challenging Enterprise Slot Filling benchmark consisting of 810 multi-turn samples and 6530 slots over 8 unique domains. Curated using a taxonomy of the 57 most challenging slot-filling scenarios observed during real-world enterprise deployments, ESF-Bench exposes notable limitations in current state-of-the-art LLMs, with GPT-OSS-120b low successfully extracting slots for only 20.7% of benchmark samples. To support continued research in this area, we publicly release the benchmark dataset, taxonomy, and accompanying evaluation code on GitHub.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。