评测视觉语言模型在真实游戏中的多模态决策与主动求知能力
StarBench: A Turn-Based RPG Benchmark for Agentic Multimodal Decision-Making and Information Seeking
- 构建基于《崩坏:星穹铁道》的回合制角色扮演游戏评测基准
- 直接控制下模型动作准确率仅32.1%,信息求助可提升成功率至68.7%
- 首次量化评估模型何时该问、如何问,适合研究智能体自主性的人
人类玩家不仅按键操作,还会将屏幕所见转化为精确的键盘鼠标动作,并在卡顿时主动寻求信息再尝试。当前视觉语言模型(VLMs)能否做到类似行为?尽管在简化控制或工具辅助下表现良好,但在真实客户端中——从原始截图映射到时序连贯的低层动作,并自主决定何时请求指导——仍是未解难题。我们提出StarBench,一个基于《崩坏:星穹铁道》的回合制角色扮演游戏基准,聚焦两大类人类能力:从像素到动作的多模态决策,以及智能体式的信息寻求。该基准统一评估八项战斗任务和两种模式(直接控制与工具辅助),使用共享任务与指标。直接控制模式下,智能体仅接收截图,需输出点击和按键等底层动作,无语义提示;工具辅助模式则通过检测器和OCR提供高层意图映射与文本化观察,降低界面理解难度。为模拟人类行为,还引入“问或做”诊断机制,衡量智能体请求帮助的时机及其对后续表现的影响。报告了当前主流VLM的基线结果与人类参考表现。结果显示,在直接控制模式下感知到动作的保真度存在显著差距(成功率为32.1%),而合理的信息寻求能显著提升性能(最高达68.7%),验证了StarBench作为可复现的评估标准,适用于真实客户端中智能体信息寻求与多模态决策的研究。
原文摘要 · Abstract (English)
Human players do more than press buttons: they ground what they see on screen into precise keyboard-mouse actions and, when stuck, they seek information before trying again. We ask whether current vision-language models (VLMs) can do the same. Despite encouraging results under simplified control or tool scaffolds, human-like play in a real client - mapping raw screenshots to temporally coherent low-level actions while deciding when to ask for guidance - remains an open challenge. We introduce StarBench, a turn-based RPG benchmark derived from Honkai: Star Rail that targets these two human-like competencies: multimodal decision-making from pixels to actions and agentic information seeking. StarBench standardizes evaluation across eight combat tasks and two regimes with shared tasks and metrics: (i) direct control, where agents receive only screenshots and must emit low-level primitives (click and keypress) with no semantic hints; and (ii) tool-assisted control, where higher-level intents can be mapped to primitives by detectors and OCR outputs provide optional textualized observations to ease UI grounding. To mirror human practice, StarBench also includes an ask-or-act diagnostic that measures whether and when agents choose to request brief guidance before proceeding, and how that choice affects subsequent performance. We report reference baselines for contemporary VLMs and a human reference. Results expose sizable gaps in perception-to-control fidelity in the direct regime, while showing that judicious information seeking correlates with improved success, establishing StarBench as a reproducible yardstick for agentic information seeking and multimodal decision-making in real-client play.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。