将音频融入大模型,让机器听懂声音并像人一样对话。
Towards General Auditory Intelligence: Large Multimodal Models for Machine Listening and Speaking
- 用大模型提升音频理解与生成能力,实现深层语义感知。
- 融合音视频信息,增强系统对场景的全面理解与推理。
- 适合关注语音交互、多模态智能的科研与工程人员。
在大语言模型(LLM)与通用人工智能(AGI)时代,计算机听觉需突破传统范式,充分借助基础模型能力,迈向更全面的理解、更自然的生成和更类人的交互。音频作为蕴含语义、情感与上下文线索的重要模态,在实现自然化、具身化的机器智能中发挥关键作用。本文综述了近期将音频融入大语言模型的进展,重点聚焦四个方向:音频理解、音频生成、基于语音的交互以及音视频联合理解。我们分析了大模型如何重塑音频感知与推理,使系统能够以更深层次的语义理解声音,生成富有表现力的音频输出,并开展类人语音交互。此外,探讨了音视频模态融合如何提升情境感知与跨模态推理能力,推动多模态智能边界。本综述不仅整合现有研究成果,还指出了构建以声音为核心的通用人工智能系统所面临的关键挑战与未来方向,使其能像人类一样通过声音感知、理解与互动。
原文摘要 · Abstract (English)
In the era of large language models (LLMs) and artificial general intelligence (AGI), computer audition must evolve beyond traditional paradigms to fully leverage the capabilities of foundation models, towards more comprehensive understanding, more natural generation and more human-like interaction. Audio, as a modality rich in semantic, emotional, and contextual cues, plays a vital role in achieving naturalistic and embodied machine intelligence. This survey provides a comprehensive review of recent progress in integrating audio into LLMs, with a focus on four key areas: audio comprehension, audio generation, speech-based interaction, and audio-visual understanding. We analyze how LLMs are reshaping audio perception and reasoning, enabling systems to understand sound at a deeper semantic level, generate expressive audio outputs, and engage in human-like spoken interaction. Furthermore, we explore how the fusion of audio and visual modalities enhances situational awareness and cross-modal reasoning, pushing the boundaries of multimodal intelligence. This survey not only synthesizes existing research but also identifies critical challenges and future directions for building audio-native AGI systems capable of perceiving, understanding, and interacting through sound as naturally as humans do.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。