系统梳理大模型时代的视听智能,打通听觉与视觉的融合研究框架。
Audio-Visual Intelligence in Large Foundation Models

- 构建统一分类体系,覆盖理解、生成与交互三类视听任务。
- 归纳跨模态融合、自回归与扩散生成等核心技术方法。
- 适合关注多模态大模型的科研人员与工业界开发者参考。
视听智能(AVI)已成为人工智能的核心前沿,连接听觉与视觉模态,使机器能在多模态真实世界中感知、生成并交互。在大模型时代,音频与视觉的联合建模愈发关键,不仅用于理解,还支持对动态时序信号的可控生成与推理。近期如Meta MovieGen和Google Veo-3等进展凸显了业界与学界对统一视听架构的关注,这些架构从海量多模态数据中学习。然而,当前研究仍碎片化,涵盖任务多样、分类标准不一、评估方式异质,阻碍系统性比较与知识整合。本文首次从大模型视角全面综述视听智能,建立涵盖理解(如语音识别、声音定位)、生成(如音驱动视频合成、视频转音频)与交互(如对话、具身或代理接口)的统一分类体系。我们梳理方法基础,包括模态标记化、跨模态融合、自回归与扩散生成、大规模预训练、指令对齐与偏好优化。同时整理代表性数据集、基准测试与评估指标,实现任务家族间的结构化对比,并指出同步性、空间推理、可控性与安全性等开放挑战。通过将快速发展的领域整合为连贯框架,本综述旨在成为未来大规模视听智能研究的基础参考。
原文摘要 · Abstract (English)
Audio-Visual Intelligence (AVI) has emerged as a central frontier in artificial intelligence, bridging auditory and visual modalities to enable machines that can perceive, generate, and interact in the multimodal real world. In the era of large foundation models, joint modeling of audio and vision has become increasingly crucial, i.e., not only for understanding but also for controllable generation and reasoning across dynamic, temporally grounded signals. Recent advances, such as Meta MovieGen and Google Veo-3, highlight the growing industrial and academic focus on unified audio-vision architectures that learn from massive multimodal data. However, despite rapid progress, the literature remains fragmented, spanning diverse tasks, inconsistent taxonomies, and heterogeneous evaluation practices that impede systematic comparison and knowledge integration. This survey provides the first comprehensive review of AVI through the lens of large foundation models. We establish a unified taxonomy covering the broad landscape of AVI tasks, ranging from understanding (e.g., speech recognition, sound localization) to generation (e.g., audio-driven video synthesis, video-to-audio) and interaction (e.g., dialogue, embodied, or agentic interfaces). We synthesize methodological foundations, including modality tokenization, cross-modal fusion, autoregressive and diffusion-based generation, large-scale pretraining, instruction alignment, and preference optimization. Furthermore, we curate representative datasets, benchmarks, and evaluation metrics, offering a structured comparison across task families and identifying open challenges in synchronization, spatial reasoning, controllability, and safety. By consolidating this rapidly expanding field into a coherent framework, this survey aims to serve as a foundational reference for future research on large-scale AVI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。