多模态视觉输入让AI在游戏里自主建结构,突破纯文本限制。
VoyagerVision: Investigating the Role of Multi-modal Information for Open-ended Learning Systems
- 用截图作为视觉反馈,让模型理解空间环境
- 50轮中平均建出2.75个新结构,远超原版无法实现
- 适合对开放世界智能体、多模态学习感兴趣的研究者
开放性是追求通用人工智能(AGI)的关键方向,使模型能自主选择任务。近期大语言模型(如GPT-4o)已具备图像理解能力,类似OMNI-EPIC的系统将视角像素数据输入模型,帮助其解析环境并完成任务。本文提出,提供视觉输入可增强模型对空间环境的理解,从而提升其开放性潜力。为此,本文提出VoyagerVision——一个基于Voyager的多模态模型,利用截图作为视觉反馈,在Minecraft中构建结构。在50轮迭代中,VoyagerVision平均成功创建2.75个独特结构,而原版Voyager无法实现此功能,标志着全新方向的拓展。在平坦世界建筑单元测试中,成功率一半,复杂结构失败率较高。项目官网见https://esmyth-dev.github.io/VoyagerVision.github.io/
原文摘要 · Abstract (English)
Open-endedness is an active field of research in the pursuit of capable Artificial General Intelligence (AGI), allowing models to pursue tasks of their own choosing. Simultaneously, recent advancements in Large Language Models (LLMs) such as GPT-4o [9] have allowed such models to be capable of interpreting image inputs. Implementations such as OMNI-EPIC [4] have made use of such features, providing an LLM with pixel data of an agent's POV to parse the environment and allow it to solve tasks. This paper proposes that providing these visual inputs to a model gives it greater ability to interpret spatial environments, and as such, can increase the number of tasks it can successfully perform, extending its open-ended potential. To this aim, this paper proposes VoyagerVision -- a multi-modal model capable of creating structures within Minecraft using screenshots as a form of visual feedback, building on the foundation of Voyager. VoyagerVision was capable of creating an average of 2.75 unique structures within fifty iterations of the system, as Voyager was incapable of this, it is an extension in an entirely new direction. Additionally, in a set of building unit tests VoyagerVision was successful in half of all attempts in flat worlds, with most failures arising in more complex structures. Project website is available at https://esmyth-dev.github.io/VoyagerVision.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。