构建首个覆盖12类游戏的LLM代理基准,支持训练与评估
Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games
- 基于MCP协议的插件化接口,支持多游戏场景研究
- 提供跨游戏类型的专家轨迹微调数据集,提升模型游戏能力
- 含排行榜、对战场等评测框架,适合游戏AI研发者使用
大型语言模型(LLM)代理正在重塑游戏产业,使角色更智能且更符合人类偏好。然而,现有游戏基准无法满足实际需求:缺乏在多种游戏类型中对LLM能力的全面评估,缺少对复杂游戏所需代理模块的研究,以及将预训练LLM转化为游戏代理所需的微调数据集。为填补这些空白,我们提出Orak,一个涵盖12款主流视频游戏的基准,覆盖所有主要游戏类型。通过基于模型上下文协议(MCP)构建的即插即用接口,Orak支持在多样化游戏场景中系统化、可复现的代理模块研究。我们还发布了包含多类型游戏专家行为轨迹的微调数据集,使通用LLM能快速转变为有效游戏代理。Orak提供统一评估框架,包括游戏排行榜、LLM对战场及输入模态、代理策略和微调效果的消融实验,为构建多功能游戏代理奠定基础。代码与数据集见https://github.com/krafton-ai/Orak 和 https://huggingface.co/datasets/KRAFTON/Orak。
原文摘要 · Abstract (English)
Large Language Model (LLM) agents are reshaping the game industry, by enabling more intelligent and human-preferable characters. Yet, current game benchmarks fall short of practical needs: they lack evaluations of diverse LLM capabilities across various game genres, studies of agentic modules crucial for complex gameplay, and fine-tuning datasets to adapt pre-trained LLMs into gaming agents. To fill these gaps, we present Orak, a benchmark for training and evaluating LLM agents across 12 popular video games spanning all major genres. Using a plug-and-play interface built on Model Context Protocol (MCP), Orak supports systematic and reproducible studies of agentic modules in varied game scenarios. We further release a fine-tuning dataset of expert LLM gameplay trajectories covering multiple genres, turning general LLMs into effective game agents. Orak offers a united evaluation framework, including game leaderboards, LLM battle arenas, and \fix{ablation studies} of input modality, agentic strategies, and fine-tuning effects, establishing a foundation towards versatile gaming agents. Code and datasets are available at https://github.com/krafton-ai/Orak and https://huggingface.co/datasets/KRAFTON/Orak.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。