arXiv:2608.20379cs.AI2026-08中稿 · TMLR综述

系统梳理多模态智能体框架的技术演进与应用前景。

A Survey on Foundations and Frontiers of Multimodal Agentic Frameworks: Techniques and Applications

论文配图:A Survey on Foundations and Frontiers of Multimodal Agentic Frameworks: Techniques and Applications
图 1 · 摘自论文原文
  • 按感知、推理、规划等模块分析多模态融合架构
  • 揭示视觉、音频等多模态如何提升智能体决策能力
  • 适合关注AI代理、具身智能的研究者参考

大语言模型(LLMs)的发展推动了智能体研究:具备推理、规划与行动能力的系统。这些智能体以强大LLM为核心,协调感知、记忆与决策。随着大规模多模态模型(LMMs)出现,系统可处理图像、音频、视频等多种模态,显著提升实际应用潜力。然而,尽管已有基于LLM的智能体综述,多模态对智能体架构的影响尚未系统考察。本文从感知、推理、规划、记忆和行动五大核心模块出发,分析多模态在智能体中的作用,梳理从文本主导到多模态框架的演化路径。探讨委托式、晚融合与早融合等多模态集成方式,评估具象感知与多模态推理带来的新型智能行为。提出以模态为中心的分类体系,关联架构设计与能力表现。覆盖机器人、图形界面与网页导航、多媒体内容生成与编辑、长视频理解与检索等应用场景,分析性能差异及训练/推理成本、延迟、部署限制等效率-可扩展性权衡。聚焦多模态对智能体设计的影响,识别关键缺口,为构建鲁棒、通用智能系统提供路线图。

原文摘要 · Abstract (English)

Advances in large language models (LLMs) have fueled a wave of research into agency: the ability to reason, plan, and act. This effort has produced agentic frameworks that orchestrate perception, memory, and decision-making around powerful LLM backbones. With the advent of large multimodal models (LMMs), these systems can process and integrate diverse modalities, including images, audio, and video, thereby improving their real-world applicability. Yet, while surveys of LLM-based agents exist, the role of multimodality in shaping agency has not been systematically examined in recent years. This survey fills the gap by analyzing the impact of multimodality across the core functional modules of the agentic framework: perception, reasoning, planning, memory, and action. Using this lens, we trace the evolution from text-centric agents to multimodal frameworks, examine how modalities are integrated through delegated, late-fusion, and early-fusion architectures, and assess the emergence of agentic behaviors enabled by grounded perception and multimodal reasoning. We organize existing work through a modality-centric taxonomy that links architectural design choices to agent capabilities. Moreover, we review multimodal agentic systems across various application domains, including Robotics, GUI & Web Navigation, Multimedia Content Generation & Editing, and Long-form Video Understanding & Retrieval. Beyond capabilities, we analyze performance across these settings and discuss efficiency-scalability trade-offs, including training and inference costs, latency, and deployment constraints. By focusing on the impact of multimodality in agentic design, we aim to identify key gaps and chart a roadmap toward robust and general-purpose intelligent systems.

智能体多模态大模型系统综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。