首篇系统综述音频语言模型,梳理语音音乐声音全场景应用
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey
- 构建统一分类体系,涵盖模型架构与训练目标
- 覆盖语音、音乐、声音三类音频任务,全面梳理研究进展
- 揭示各研究方向的互动与制约,指引未来技术演进
音频语言模型(ALMs)通过配对的音视频数据训练,旨在处理、理解与推理以音频为核心的多模态内容。与依赖预定义标签的传统监督方法不同,ALMs利用自然语言作为监督信号,更适用于包含多重重叠事件的复杂真实音频场景。尽管展现出卓越的零样本与任务泛化能力,现有研究仍缺乏系统性综述来全面组织与分析发展脉络。本文首次提出对ALMs的系统性回顾,主要贡献包括:(1) 从通用音频视角,全面覆盖语音、音乐与声音领域的ALM研究;(2) 建立统一的ALM基础分类体系,涵盖模型架构与训练目标;(3) 构建研究格局,揭示不同研究维度间的相互促进与制约关系,有助于总结评估、局限性及未来方向。本综述为研究人员理解现有技术演进与未来趋势提供支持,并为实际应用实现提供参考。
原文摘要 · Abstract (English)
Audio-Language Models (ALMs), trained on paired audio-text data, are designed to process, understand, and reason about audio-centric multimodal content. Unlike traditional supervised approaches that use predefined labels, ALMs leverage natural language supervision to better handle complex real-world audio scenes with multiple overlapping events. While demonstrating impressive zero-shot and task generalization capabilities, there is still a notable lack of systematic surveys that comprehensively organize and analyze developments. In this paper, we present the first systematic review of ALMs with three main contributions: (1) comprehensive coverage of ALM works across speech, music, and sound from a general audio perspective; (2) a unified taxonomy of ALM foundations, including model architectures and training objectives; (3) establishment of a research landscape capturing mutual promotion and constraints among different research aspects, aiding in summarizing evaluations, limitations, concerns and promising directions. Our review contributes to helping researchers understand the development of existing technologies and future trends, while also providing valuable references for implementation in practical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。