用合成环境模拟人类学习语言与视觉概念的组合能力
Human-like compositional learning of visually-grounded concepts using synthetic environments
- 通过强化学习在3D合成环境中理解带限定词和介词的语言指令
- 课程学习使训练次数减少15%,让模型掌握复杂空间关系
- 能快速适应未见过的物体组合,展现类人泛化能力
语言的组合结构使人类能够分解复杂短语并映射到新视觉概念,展现灵活智能。尽管一些算法表现出组合性,却未能揭示人类如何通过试错学习组合概念并将其与视觉线索关联。为此,我们设计了一个3D合成环境,让智能体通过强化学习根据自然语言指令导航至目标。指令包含名词、属性,以及关键的限定词、介词或两者兼具。大量词汇组合提升了视觉定位任务的组合复杂度——例如,指令为“上方红色球体下方的蓝色立方体”时,若指令是“某些蓝色立方体在红球下方”,则导航至蓝色立方体上方的红色球体不会被奖励。我们首先证明,强化学习智能体可将限定词概念与视觉目标对齐,但难以掌握复杂的介词概念。其次,我们发现课程学习(人类采用的学习策略)显著提升学习效率,在限定词环境中减少15%训练轮次,并使智能体轻松学会介词概念。最后,我们证实,经限定词或介词训练的智能体能分解未见测试指令,并快速适应未知物体组合的导航策略。借助合成环境,我们的研究展示了多模态强化学习智能体可实现对复杂概念类别的组合理解,并凸显类人学习策略在提升人工智能系统学习效率方面的有效性。
原文摘要 · Abstract (English)
The compositional structure of language enables humans to decompose complex phrases and map them to novel visual concepts, showcasing flexible intelligence. While several algorithms exhibit compositionality, they fail to elucidate how humans learn to compose concept classes and ground visual cues through trial and error. To investigate this multi-modal learning challenge, we designed a 3D synthetic environment in which an agent learns, via reinforcement, to navigate to a target specified by a natural language instruction. These instructions comprise nouns, attributes, and critically, determiners, prepositions, or both. The vast array of word combinations heightens the compositional complexity of the visual grounding task, as navigating to a blue cube above red spheres is not rewarded when the instruction specifies navigating to "some blue cubes below the red sphere". We first demonstrate that reinforcement learning agents can ground determiner concepts to visual targets but struggle with more complex prepositional concepts. Second, we show that curriculum learning, a strategy humans employ, enhances concept learning efficiency, reducing the required training episodes by 15% in determiner environments and enabling agents to easily learn prepositional concepts. Finally, we establish that agents trained on determiner or prepositional concepts can decompose held-out test instructions and rapidly adapt their navigation policies to unseen visual object combinations. Leveraging synthetic environments, our findings demonstrate that multi-modal reinforcement learning agents can achieve compositional understanding of complex concept classes and highlight the efficacy of human-like learning strategies in improving artificial systems' learning efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。