让游戏机器人听懂声音指令,实现多模态交互
STEVE-Audio: Expanding the Goal Conditioning Modalities of Embodied Agents in Minecraft
- 将音频输入映射到代理的潜在目标空间,扩展控制方式
- 音频条件代理表现与文本/视觉条件相当,成功率接近基准
- 开源代码与模型,推动多模态决策研究
近期,STEVE-1方法被提出用于训练生成式代理,使其能够根据潜空间中的CLIP嵌入执行指令。本文提出一种新方法,通过学习将新输入模态映射至代理的潜空间目标,扩展控制方式。我们在具有挑战性的Minecraft环境中应用该方法,将目标条件扩展至音频模态。所提出的音频条件代理在性能上可与原始文本和视觉条件代理相媲美。具体而言,我们构建了适用于Minecraft的音视频CLIP基础模型及音频先验网络,共同将音频样本映射到STEVE-1策略的潜空间目标。此外,我们还分析了不同模态条件带来的权衡。训练代码、评估代码及Minecraft音视频CLIP基础模型已开源,以促进多模态通用序列决策代理的进一步研究。
原文摘要 · Abstract (English)
Recently, the STEVE-1 approach has been introduced as a method for training generative agents to follow instructions in the form of latent CLIP embeddings. In this work, we present a methodology to extend the control modalities by learning a mapping from new input modalities to the latent goal space of the agent. We apply our approach to the challenging Minecraft domain, and extend the goal conditioning to include the audio modality. The resulting audio-conditioned agent is able to perform on a comparable level to the original text-conditioned and visual-conditioned agents. Specifically, we create an Audio-Video CLIP foundation model for Minecraft and an audio prior network which together map audio samples to the latent goal space of the STEVE-1 policy. Additionally, we highlight the tradeoffs that occur when conditioning on different modalities. Our training code, evaluation code, and Audio-Video CLIP foundation model for Minecraft are made open-source to help foster further research into multi-modal generalist sequential decision-making agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。