评估大模型生成流程模型的能力,发现其结果是多组权衡方案而非单一最优解。
Evaluating the Process Modeling Abilities of Large Language Models -- Preliminary Foundations and Results
- 大模型生成的不是唯一最优流程模型,而是多个权衡后的可行方案。
- 评估需综合考虑模型质量、生成耗时与成本,不能只看输出结果。
- 适合关注AI生成流程建模可信度的研究者与系统设计人员参考。
大语言模型(LLM)已彻底改变自然语言处理领域。尽管初步测试显示其在流程建模方面潜力可观,但目前仍存在争议:大模型究竟能在多大程度上生成高质量的流程模型。本文指出,评估大模型的流程建模能力远非简单任务。例如,在一个简单场景中,不仅应评估模型质量,还需考量生成所需的时间与成本。因此,大模型并非生成单一最优解,而是产出一组帕累托最优的备选方案。此外,还存在诸多挑战,如质量概念的界定、结果验证、泛化能力及数据泄露问题。本文深入讨论这些挑战,并提出未来可开展的科学实验以系统应对。
原文摘要 · Abstract (English)
Large language models (LLM) have revolutionized the processing of natural language. Although first benchmarks of the process modeling abilities of LLM are promising, it is currently under debate to what extent an LLM can generate good process models. In this contribution, we argue that the evaluation of the process modeling abilities of LLM is far from being trivial. Hence, available evaluation results must be taken carefully. For example, even in a simple scenario, not only the quality of a model should be taken into account, but also the costs and time needed for generation. Thus, an LLM does not generate one optimal solution, but a set of Pareto-optimal variants. Moreover, there are several further challenges which have to be taken into account, e.g. conceptualization of quality, validation of results, generalizability, and data leakage. We discuss these challenges in detail and discuss future experiments to tackle these challenges scientifically.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。