arXiv:2501.03151cs.AIcs.CV2025-01综述被引 23

剖析大模型实现通用智能的四大基础难题与突破路径

Large language models for artificial general intelligence (AGI): A survey of foundational principles and approaches

  • 从具身、符号接地、因果和记忆四方面重构大模型认知能力
  • 指出当前大模型虽能对话推理但本质仍浅层脆弱,难达通用智能
  • 适合关注大模型局限与通用人工智能理论框架的研究者阅读

基于大规模预训练基础模型(PFMs)的生成式人工智能系统,如视觉语言模型、大语言模型(LLMs)、扩散模型和视觉-语言-动作(VLA)模型,在多种领域和情境中展现出解决复杂非平凡问题的能力。多模态大语言模型(MLLMs)通过海量多样化数据学习,形成对世界的丰富细腻表征,具备推理、有意义对话、人机协同解决复杂问题以及理解人类社会情感等广泛能力。然而,最先进的大模型在大规模数据上训练后,其认知能力仍显表面化且脆弱,通用能力严重受限。实现人类水平的通用智能(AGI)需解决具身性、符号接地、因果关系和记忆等基础性问题。这些概念更贴近人类认知,可为大模型赋予内在的人类认知特性,支持物理合理性、语义明确性、灵活性和更强泛化性的知识与智能。本文探讨上述基础问题,并综述将这些原则有机融入大模型以推动实现AGI的前沿方法。

原文摘要 · Abstract (English)

Generative artificial intelligence (AI) systems based on large-scale pretrained foundation models (PFMs) such as vision-language models, large language models (LLMs), diffusion models and vision-language-action (VLA) models have demonstrated the ability to solve complex and truly non-trivial AI problems in a wide variety of domains and contexts. Multimodal large language models (MLLMs), in particular, learn from vast and diverse data sources, allowing rich and nuanced representations of the world and, thereby, providing extensive capabilities, including the ability to reason, engage in meaningful dialog; collaborate with humans and other agents to jointly solve complex problems; and understand social and emotional aspects of humans. Despite this impressive feat, the cognitive abilities of state-of-the-art LLMs trained on large-scale datasets are still superficial and brittle. Consequently, generic LLMs are severely limited in their generalist capabilities. A number of foundational problems -- embodiment, symbol grounding, causality and memory -- are required to be addressed for LLMs to attain human-level general intelligence. These concepts are more aligned with human cognition and provide LLMs with inherent human-like cognitive properties that support the realization of physically-plausible, semantically meaningful, flexible and more generalizable knowledge and intelligence. In this work, we discuss the aforementioned foundational issues and survey state-of-the art approaches for implementing these concepts in LLMs. Specifically, we discuss how the principles of embodiment, symbol grounding, causality and memory can be leveraged toward the attainment of artificial general intelligence (AGI) in an organic manner.

大模型通用智能认知架构符号接地

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。