arXiv:2604.01929cs.SDcs.AI2026-04被引 3

Woosh是索尼推出的音效生成基础模型,支持文本和视频生成高质量音效。

Woosh: A Sound Effects Foundation Model

论文配图:Woosh: A Sound Effects Foundation Model
图 1 · 摘自论文原文
  • 构建音效专用编码器、对齐模型及文本/视频到音效生成模块
  • 在公开与私有数据上表现优于或媲美StableAudio-Open等开源模型
  • 提供轻量化模型,支持低资源环境快速推理,适合音频开发与创作

音频研究社区依赖开源生成模型作为构建新方法和建立基线的基础工具。本文介绍Sony AI公开发布的音效基础模型Woosh,详述其架构、训练过程,并与现有主流开源模型进行评估对比。针对音效优化,我们提供了(1)高质量音频编码器/解码器模型,(2)用于条件控制的文本-音频对齐模型,以及(3)文本到音频、(4)视频到音频生成模型。释放版本还包含蒸馏后的文本到音频和视频到音频模型,支持低资源运行与快速推理。在公共与私有数据上的评估表明,各模块性能均达到或超过StableAudio-Open、TangoFlux等现有开源模型。推理代码与模型权重已公开于https://github.com/SonyResearch/Woosh,演示样本见https://sonyresearch.github.io/Woosh/。

原文摘要 · Abstract (English)

The audio research community depends on open generative models as foundational tools for building novel approaches and establishing baselines. In this report, we present Woosh, Sony AI's publicly released sound effect foundation model, detailing its architecture, training process, and an evaluation against other popular open models. Being optimized for sound effects, we provide (1) a high-quality audio encoder/decoder model and (2) a text-audio alignment model for conditioning, together with (3) text-to-audio and (4) video-to-audio generative models. Distilled text-to-audio and video-to-audio models are also included in the release, allowing for low-resource operation and fast inference. Our evaluation on both public and private data shows competitive or better performance for each module when compared to existing open alternatives like StableAudio-Open and TangoFlux. Inference code and model weights are available at https://github.com/SonyResearch/Woosh. Demo samples can be found at https://sonyresearch.github.io/Woosh/.

音效生成基础模型文本到音频视频到音频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。