arXiv:2505.14552cs.CLcs.AI2025-05

KORGym用动态游戏评估大模型推理能力,更全面真实。

KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation

  • 设计50多个文本/视觉互动游戏,支持多轮强化学习测试。
  • 19个大模型和8个视觉模型实验显示闭源模型表现更优。
  • 适合研究复杂交互环境下模型推理的学者使用。

大语言模型(LLMs)的发展迫切需要更全面的评估方法以准确衡量其推理能力。现有基准多局限于特定领域,难以反映模型的通用推理潜力。为此,我们提出知识正交推理体操馆(KORGym),受KOR-Bench和Gymnasium启发,提供超过五十个文本或视觉形式的游戏,并支持与强化学习场景结合的交互式多轮评估。在19个LLMs和8个VLMs上进行大规模实验,揭示了模型家族内一致的推理模式,且表明闭源模型具有更优性能。进一步分析了模态、推理策略、强化学习技术及响应长度对模型表现的影响。我们期望KORGym能成为推动大模型推理研究和复杂交互环境评估方法发展的关键资源。

原文摘要 · Abstract (English)

Recent advancements in large language models (LLMs) underscore the need for more comprehensive evaluation methods to accurately assess their reasoning capabilities. Existing benchmarks are often domain-specific and thus cannot fully capture an LLM's general reasoning potential. To address this limitation, we introduce the Knowledge Orthogonal Reasoning Gymnasium (KORGym), a dynamic evaluation platform inspired by KOR-Bench and Gymnasium. KORGym offers over fifty games in either textual or visual formats and supports interactive, multi-turn assessments with reinforcement learning scenarios. Using KORGym, we conduct extensive experiments on 19 LLMs and 8 VLMs, revealing consistent reasoning patterns within model families and demonstrating the superior performance of closed-source models. Further analysis examines the effects of modality, reasoning strategies, reinforcement learning techniques, and response length on model performance. We expect KORGym to become a valuable resource for advancing LLM reasoning research and developing evaluation methodologies suited to complex, interactive environments.

大模型评估推理能力强化学习动态游戏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。