arXiv:2410.02365cs.CLcs.AI2024-10被引 3

用视觉和语言融合学习高阶抽象概念,让模型像人一样理解抽象概念。

From Concrete to Abstract: A Multimodal Generative Approach to Abstract Concept Learning

  • 通过视觉与语言信息融合,从具体概念逐步抽象到高层概念。
  • 在语言理解与命名任务中表现优异,验证了抽象学习能力。
  • 适合研究认知建模、多模态学习的学者参考。

理解与操控具体与抽象概念是人类智能的核心。然而,这对人工系统仍具挑战。本文提出一种多模态生成方法,用于高阶抽象概念学习,整合具体概念的视觉与类别语言信息。模型首先将下位概念具象化,合并形成基本层次概念,再基于基本层次概念的具象化,抽象出上位概念。通过语言到视觉与视觉到语言测试评估模型的语言学习能力,实验结果表明模型在语言理解与命名任务中均表现出色。

原文摘要 · Abstract (English)

Understanding and manipulating concrete and abstract concepts is fundamental to human intelligence. Yet, they remain challenging for artificial agents. This paper introduces a multimodal generative approach to high order abstract concept learning, which integrates visual and categorical linguistic information from concrete ones. Our model initially grounds subordinate level concrete concepts, combines them to form basic level concepts, and finally abstracts to superordinate level concepts via the grounding of basic-level concepts. We evaluate the model language learning ability through language-to-visual and visual-to-language tests with high order abstract concepts. Experimental results demonstrate the proficiency of the model in both language understanding and language naming tasks.

抽象概念多模态生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。