构建过程挖掘场景下的LLM计划生成数据集,评估复杂任务规划能力。
ProcessTBench: An LLM Plan Generation Dataset for Process Mining
- 基于任务挖掘框架设计合成数据集,支持多语言与并行动作
- 包含2000个带语义变体的任务指令,覆盖12种流程模式
- 适合研究LLM在真实流程场景中的推理与执行一致性
大型语言模型(LLMs)在计划生成方面展现出显著潜力。然而,现有数据集往往缺乏高级工具使用场景所需的复杂性——例如处理改写查询、支持多语言以及管理可并行执行的动作。这些场景对评估LLMs在现实应用中的演进能力至关重要。此外,当前数据集无法从流程视角研究LLM的表现,尤其在理解同一流程在不同条件或表述下的典型行为与挑战方面存在不足。为此,我们提出了ProcessTBench合成数据集,作为TaskBench的扩展,专为在流程挖掘框架下评估LLMs而设计。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have shown significant promise in plan generation. Yet, existing datasets often lack the complexity needed for advanced tool use scenarios - such as handling paraphrased query statements, supporting multiple languages, and managing actions that can be done in parallel. These scenarios are crucial for evaluating the evolving capabilities of LLMs in real-world applications. Moreover, current datasets don't enable the study of LLMs from a process perspective, particularly in scenarios where understanding typical behaviors and challenges in executing the same process under different conditions or formulations is crucial. To address these gaps, we present the ProcessTBench synthetic dataset, an extension of the TaskBench dataset specifically designed to evaluate LLMs within a process mining framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。