arXiv:2410.03255cs.AIcs.CL2024-10被引 7

为业务流程管理任务构建首个专用大模型评测基准。

Towards a Benchmark for Large Language Models for Business Process Management Tasks

  • 设计四类BPM任务,对比开源小模型表现
  • 发现模型大小与任务类型显著影响性能差异
  • 指导企业根据需求选型,尤其关注开源模型

越来越多组织将大语言模型(LLMs)应用于各类任务。尽管通用性强,但LLMs常出现错误甚至幻觉。现有性能评测往往难以映射到具体实际场景。本文填补了业务流程管理(BPM)领域缺乏专用评测基准的空白。通过系统比较四类BPM任务中多个小型开源模型的表现,分析任务特异性、开源与商用模型效果差异,以及模型规模对性能的影响。研究结果为组织在实际应用中选择合适模型提供依据。

原文摘要 · Abstract (English)

An increasing number of organizations are deploying Large Language Models (LLMs) for a wide range of tasks. Despite their general utility, LLMs are prone to errors, ranging from inaccuracies to hallucinations. To objectively assess the capabilities of existing LLMs, performance benchmarks are conducted. However, these benchmarks often do not translate to more specific real-world tasks. This paper addresses the gap in benchmarking LLM performance in the Business Process Management (BPM) domain. Currently, no BPM-specific benchmarks exist, creating uncertainty about the suitability of different LLMs for BPM tasks. This paper systematically compares LLM performance on four BPM tasks focusing on small open-source models. The analysis aims to identify task-specific performance variations, compare the effectiveness of open-source versus commercial models, and assess the impact of model size on BPM task performance. This paper provides insights into the practical applications of LLMs in BPM, guiding organizations in selecting appropriate models for their specific needs.

大模型评测业务流程开源模型LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。