arXiv:2507.00951cs.AI2025-07被引 12

突破纯文本预测,探索具身智能与认知协同的AGI新路径

Thinking Beyond Tokens: From Brain-Inspired Intelligence to Cognitive Foundations for Artificial General Intelligence and its Societal Impact

  • 以脑科学为灵感,构建模块化推理与持续记忆融合的智能架构
  • 提出动态工具调用与多智能体协作框架,提升系统适应性
  • 适合关注AGI本质、认知模型与人机协同的研究者

机器能否像人类一样思考、推理和行动?尽管GPT-4.5、DeepSeek、Claude 3.5 Sonnet、Phi-4和Grok 3等模型展现出多模态流畅性和部分推理能力,仍受限于基于标记(token)的预测范式,缺乏具身代理性。本文跨人工智能、认知神经科学、心理学、生成模型与智能体系统,分析通用智能的架构与认知基础,强调模块化推理、持久记忆与多智能体协同的重要性。重点探讨了结合检索、规划与动态工具使用的代理型RAG框架,推动更自适应行为。讨论了信息压缩、测试时适应与训练零成本方法等泛化策略,作为迈向领域无关智能的关键路径。视觉-语言模型被重新审视,不再仅作感知模块,而是演变为具身理解与协作任务完成的接口。我们主张,真正的智能并非来自规模本身,而是记忆与推理的集成:模块化、交互式、自我优化组件的协调,其中压缩机制支撑适应行为。结合神经符号系统、强化学习与认知支架进展,探索统计学习与目标导向认知间的桥梁。最后,识别通往AGI的关键科学、技术和伦理挑战。

原文摘要 · Abstract (English)

Can machines truly think, reason and act in domains like humans? This enduring question continues to shape the pursuit of Artificial General Intelligence (AGI). Despite the growing capabilities of models such as GPT-4.5, DeepSeek, Claude 3.5 Sonnet, Phi-4, and Grok 3, which exhibit multimodal fluency and partial reasoning, these systems remain fundamentally limited by their reliance on token-level prediction and lack of grounded agency. This paper offers a cross-disciplinary synthesis of AGI development, spanning artificial intelligence, cognitive neuroscience, psychology, generative models, and agent-based systems. We analyze the architectural and cognitive foundations of general intelligence, highlighting the role of modular reasoning, persistent memory, and multi-agent coordination. In particular, we emphasize the rise of Agentic RAG frameworks that combine retrieval, planning, and dynamic tool use to enable more adaptive behavior. We discuss generalization strategies, including information compression, test-time adaptation, and training-free methods, as critical pathways toward flexible, domain-agnostic intelligence. Vision-Language Models (VLMs) are reexamined not just as perception modules but as evolving interfaces for embodied understanding and collaborative task completion. We also argue that true intelligence arises not from scale alone but from the integration of memory and reasoning: an orchestration of modular, interactive, and self-improving components where compression enables adaptive behavior. Drawing on advances in neurosymbolic systems, reinforcement learning, and cognitive scaffolding, we explore how recent architectures begin to bridge the gap between statistical learning and goal-directed cognition. Finally, we identify key scientific, technical, and ethical challenges on the path to AGI.

AGI认知模型智能体神经符号

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。