用扩散模型实现更快更丰富的音频描述生成。
Towards Diverse and Efficient Audio Captioning via Diffusion Models
- 采用非自回归扩散模型,整体建模音频上下文。
- 生成速度和多样性均优于现有方法,质量达顶尖水平。
- 适合需要高效多样文本生成的多模态应用。
我们提出基于扩散模型的音频描述生成(DAC),一种面向多样化与高效性音频描述的任务框架。尽管依赖语言主干的现有描述模型在多个任务中取得显著进展,但其生成速度慢、多样性不足的问题限制了音频理解与多媒体应用的发展。我们的扩散模型框架利用其固有的随机性和全局上下文建模能力,展现出独特优势。通过严格评估,DAC不仅在描述质量上达到当前最佳性能,且在生成速度与多样性方面显著超越已有方法。结果表明,文本生成可无缝集成于以扩散模型为基础的音频与视觉生成任务中,为跨模态统一生成模型提供了新路径。
原文摘要 · Abstract (English)
We introduce Diffusion-based Audio Captioning (DAC), a non-autoregressive diffusion model tailored for diverse and efficient audio captioning. Although existing captioning models relying on language backbones have achieved remarkable success in various captioning tasks, their insufficient performance in terms of generation speed and diversity impede progress in audio understanding and multimedia applications. Our diffusion-based framework offers unique advantages stemming from its inherent stochasticity and holistic context modeling in captioning. Through rigorous evaluation, we demonstrate that DAC not only achieves SOTA performance levels compared to existing benchmarks in the caption quality, but also significantly outperforms them in terms of generation speed and diversity. The success of DAC illustrates that text generation can also be seamlessly integrated with audio and visual generation tasks using a diffusion backbone, paving the way for a unified, audio-related generative model across different modalities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。