arXiv:2510.08928cs.AI2025-10被引 5

用打斗游戏测试大模型实时决策能力,更真实地评估AI策略思维。

LM Fight Arena: Benchmarking Large Multimodal Models via Game Competition

  • 让大模型在《真人快打2》中互相对战,通过实时画面理解选招
  • 6个主流模型在相同角色下对战,胜率差异反映真实策略水平
  • 适合关注AI交互能力、游戏智能的开发者与研究者

现有大模型多模态评估常无法反映其在实时对抗环境中的表现。本文提出LM Fight Arena,一个基于经典格斗游戏《真人快打2》的新评测框架,要求模型实时解析游戏画面与状态数据,做出战术决策。在受控锦标赛中,六款主流开源与闭源模型以相同角色对战,确保公平性。该框架实现全自动、可复现、客观的动态评估,全面检验模型的战略推理能力。本工作构建了一个兼具挑战性与趣味性的评测体系,弥合了AI评估与互动娱乐之间的鸿沟。

原文摘要 · Abstract (English)

Existing benchmarks for large multimodal models (LMMs) often fail to capture their performance in real-time, adversarial environments. We introduce LM Fight Arena (Large Model Fight Arena), a novel framework that evaluates LMMs by pitting them against each other in the classic fighting game Mortal Kombat II, a task requiring rapid visual understanding and tactical, sequential decision-making. In a controlled tournament, we test six leading open- and closed-source models, where each agent operates controlling the same character to ensure a fair comparison. The models are prompted to interpret game frames and state data to select their next actions. Unlike static evaluations, LM Fight Arena provides a fully automated, reproducible, and objective assessment of an LMM's strategic reasoning capabilities in a dynamic setting. This work introduces a challenging and engaging benchmark that bridges the gap between AI evaluation and interactive entertainment.

多模态评估游戏智能策略推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。