用模拟任务评测大模型创造力,发现闭源模型仍领先开源模型。
SimulBench: Evaluating Language Models with Creative Simulation Tasks
- 用固定LLM作用户代理生成多轮交互数据
- 在18.55%场景中GPT-4-turbo优于LLaMA-3-70b-Chat
- 适合评估模型在真实交互中的应变能力
我们提出SimulBench,一个涵盖多种创意模拟场景的基准,如扮演Linux终端或与用户玩文字游戏。这些模拟任务能有效衡量大语言模型的通用智能,但很少被纳入现有评测体系。核心挑战在于如何公平评估不同模型,同时保持任务的多轮交互特性。为此,我们采用固定LLM作为用户代理,先与目标模型生成对话,再提取高难度对话脚本用于评测。为实现自动评估,在DataName上使用GPT-4作为评分者,判断目标模型回复质量。全面实验表明,这些模拟任务仍具显著挑战性,且闭源模型与先进开源模型间存在差距。例如,GPT-4-turbo在18.55%更多案例中表现更优。
原文摘要 · Abstract (English)
We introduce SimulBench, a benchmark designed to evaluate large language models (LLMs) across a diverse collection of creative simulation scenarios, such as acting as a Linux terminal or playing text games with users. While these simulation tasks serve as effective measures of an LLM's general intelligence, they are seldom incorporated into existing benchmarks. A major challenge is to develop an evaluation framework for testing different LLMs fairly while preserving the multi-round interactive nature of simulation tasks between users and AI. To tackle this issue, we suggest using a fixed LLM as a user agent to engage with an LLM to collect dialogues first under different tasks. Then, challenging dialogue scripts are extracted for evaluating different target LLMs. To facilitate automatic assessment on \DataName{}, GPT-4 is employed as the evaluator, tasked with reviewing the quality of the final response generated by the target LLMs given multi-turn dialogue scripts. Our comprehensive experiments indicate that these simulation tasks continue to pose a significant challenge with their unique natures and show the gap between proprietary models and the most advanced open LLMs. For example, GPT-4-turbo outperforms LLaMA-3-70b-Chat on 18.55\% more cases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。