让声音分离支持文本、图像和时间跨度三种提示方式,实现通用音频分割。
SAM Audio: Segment Anything in Audio
- 统一文本、视觉掩码与时间跨度三种提示模态,构建通用音频分割框架。
- 在真实场景与专业音频中均超越现有通用与专用模型,性能领先。
- 引入新基准与无参考评估模型,更贴近人类听觉判断。
通用音频源分离是多模态人工智能系统感知与理解声音的关键能力。尽管近年进展显著,现有分离模型仍受限于领域特定(如仅针对语音或音乐)或可控性差,仅支持单一提示方式(如仅文本)。本文提出SAM Audio,一个统一文本、视觉与时间跨度提示的音频分割基础模型。基于扩散变换器架构,采用流匹配在大规模跨域音频数据(涵盖语音、音乐与一般声响)上训练,可灵活分离由语言、视觉掩码或时间片段描述的目标声源。在涵盖通用声音、语音、音乐及乐器分离的多样化基准测试中表现卓越,显著优于现有通用与专用系统。此外,我们构建了一个带人工标注多模态提示的真实世界分离基准,并设计了一种与人类判断高度相关的无参考评估模型。
原文摘要 · Abstract (English)
General audio source separation is a key capability for multimodal AI systems that can perceive and reason about sound. Despite substantial progress in recent years, existing separation models are either domain-specific, designed for fixed categories such as speech or music, or limited in controllability, supporting only a single prompting modality such as text. In this work, we present SAM Audio, a foundation model for general audio separation that unifies text, visual, and temporal span prompting within a single framework. Built on a diffusion transformer architecture, SAM Audio is trained with flow matching on large-scale audio data spanning speech, music, and general sounds, and can flexibly separate target sources described by language, visual masks, or temporal spans. The model achieves state-of-the-art performance across a diverse suite of benchmarks, including general sound, speech, music, and musical instrument separation in both in-the-wild and professionally produced audios, substantially outperforming prior general-purpose and specialized systems. Furthermore, we introduce a new real-world separation benchmark with human-labeled multimodal prompts and a reference-free evaluation model that correlates strongly with human judgment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。