arXiv:2507.15501cs.CLcs.AI2025-07ACL被引 1

用模拟环境评估大模型执行复杂任务的能力

ASPERA: A Simulated Environment to Evaluate Planning for Complex Action Execution

  • 构建模拟助手库与人机协同数据生成引擎
  • 释放250个高难度任务数据集,验证模型挑战
  • 适合研究智能助理与复杂指令执行的学者

本工作评估大语言模型(LLMs)在驱动数字助理执行复杂任务方面的潜力。这些助理依赖预训练编程知识,通过组合助手库中的对象与函数,生成多步骤动作执行程序。为此,我们开发了ASPERA框架,包含助手库模拟和人机协同的LLM数据生成引擎。该引擎使开发者能引导生成高质量任务,包括复杂用户查询、仿真状态及对应验证程序,解决了数据稀缺与评估鲁棒性问题。我们还发布了基于ASPERA生成的Asper-Bench数据集,包含250个具有挑战性的任务,用于证明:相较于无依赖代码生成,基于自定义助手库的程序生成对LLMs仍是重大挑战。

原文摘要 · Abstract (English)

This work evaluates the potential of large language models (LLMs) to power digital assistants capable of complex action execution. These assistants rely on pre-trained programming knowledge to execute multi-step goals by composing objects and functions defined in assistant libraries into action execution programs. To achieve this, we develop ASPERA, a framework comprising an assistant library simulation and a human-assisted LLM data generation engine. Our engine allows developers to guide LLM generation of high-quality tasks consisting of complex user queries, simulation state and corresponding validation programs, tackling data availability and evaluation robustness challenges. Alongside the framework we release Asper-Bench, an evaluation dataset of 250 challenging tasks generated using ASPERA, which we use to show that program generation grounded in custom assistant libraries is a significant challenge to LLMs compared to dependency-free code generation.

大模型智能助理任务规划数据生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。