构建179个可调控难度的视觉环境,助力智能体视觉学习研究
Gym-V: A Unified Vision Environment System for Agentic Vision Research
- 统一平台整合10大领域179个程序生成环境,支持可控实验
- 发现观察辅助比强化学习算法更影响训练成败,文本提示决定能否学习
- 多轮交互放大泛化与负迁移效应,适合研究视觉语言模型智能体
随着智能体越来越依赖可验证奖励的强化学习,标准化的'gym'基础设施已成为快速迭代、可复现性和公平比较的关键。当前视觉智能体缺乏此类基础设施,限制了对其学习驱动力和现有模型短板的系统研究。我们提出Gym-V,一个包含179个程序生成视觉环境的统一平台,覆盖10个领域且难度可调,使此前在分散工具包中无法实现的受控实验成为可能。使用该平台,我们发现观察辅助(如描述和游戏规则)对训练成功的影响远超强化学习算法的选择,甚至决定学习能否发生。跨域迁移实验进一步表明,多样化任务训练具有广泛泛化能力,而狭窄训练可能导致负迁移,多轮交互会放大这些效应。Gym-V已开源,作为训练环境与评估工具的基础,旨在加速未来视觉语言模型智能体的研究。
原文摘要 · Abstract (English)
As agentic systems increasingly rely on reinforcement learning from verifiable rewards, standardized ``gym'' infrastructure has become essential for rapid iteration, reproducibility, and fair comparison. Vision agents lack such infrastructure, limiting systematic study of what drives their learning and where current models fall short. We introduce \textbf{Gym-V}, a unified platform of 179 procedurally generated visual environments across 10 domains with controllable difficulty, enabling controlled experiments that were previously infeasible across fragmented toolkits. Using it, we find that observation scaffolding is more decisive for training success than the choice of RL algorithm, with captions and game rules determining whether learning succeeds at all. Cross-domain transfer experiments further show that training on diverse task categories generalizes broadly while narrow training can cause negative transfer, with multi-turn interaction amplifying all of these effects. Gym-V is released as a convenient foundation for training environments and evaluation toolkits, aiming to accelerate future research on agentic VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。