构建17个可定制的多模态交互环境,评测视觉语言模型长程决策能力。
VisGym: Diverse, Customizable, Scalable Environments for Multimodal Agents
- 设计17个跨任务环境,支持难度、输入和反馈灵活配置
- 前沿模型在复杂任务中成功率仅26.0%,长上下文反而表现更差
- 显式目标观测与探索性演示能显著提升模型性能,适合强化学习研究者
当前视觉语言模型在多步骤视觉交互中的表现仍不清晰,尤其在感知、记忆与动作的长期整合方面。我们提出VisGym,一个包含17个环境的评估与训练平台,涵盖符号谜题、真实图像理解、导航与操作任务,支持对难度、输入表示、规划时长和反馈机制的灵活控制。同时提供多步求解器生成结构化示范,支持监督微调。评估显示,所有前沿模型在交互设置中表现不佳,简单配置成功率为46.6%,困难配置仅为26.0%。实验揭示显著局限:模型难以有效利用长上下文,在无限制历史下表现劣于截断窗口;且文本符号任务在视觉呈现后难度显著上升。然而,在部分可观测或动态未知环境中,显式目标观测、文本反馈及探索性示范可带来稳定提升,揭示了关键失败模式与改进路径。代码、数据与模型见:https://visgym.github.io/。
原文摘要 · Abstract (English)
Modern Vision-Language Models (VLMs) remain poorly characterized in multi-step visual interactions, particularly in how they integrate perception, memory, and action over long horizons. We introduce VisGym, a gymnasium of 17 environments for evaluating and training VLMs. The suite spans symbolic puzzles, real-image understanding, navigation, and manipulation, and provides flexible controls over difficulty, input representation, planning horizon, and feedback. We also provide multi-step solvers that generate structured demonstrations, enabling supervised finetuning. Our evaluations show that all frontier models struggle in interactive settings, achieving low success rates in both the easy (46.6%) and hard (26.0%) configurations. Our experiments reveal notable limitations: models struggle to effectively leverage long context, performing worse with an unbounded history than with truncated windows. Furthermore, we find that several text-based symbolic tasks become substantially harder once rendered visually. However, explicit goal observations, textual feedback, and exploratory demonstrations in partially observable or unknown-dynamics settings for supervised finetuning yield consistent gains, highlighting concrete failure modes and pathways for improving multi-step visual decision-making. Code, data, and models can be found at: https://visgym.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。