arXiv:2502.16584cs.SDcs.AI2025-02被引 1

构建100万+实例的音频指令数据集,让大模型同时懂听会生成语音音乐和声音。

Audio-FLAN: An Instruction-Following Dataset for Unified Audio Understanding and Generation of Speech, Music, and Sound

  • 基于80类跨域任务构建统一音频指令数据集
  • 覆盖100万以上实例,支持零样本理解与生成
  • 适合研究多模态大模型、音频生成与理解的开发者

近年来,音频分词技术显著提升了音频能力在大型语言模型(LLMs)中的整合。然而,音频理解和生成常被视作独立任务,阻碍了真正统一的音视频-语言模型的发展。尽管指令微调在文本和视觉领域已展现出卓越的泛化与零样本学习能力,其在音频领域的应用仍基本空白。主要障碍在于缺乏能统一音频理解与生成的综合性数据集。为此,我们提出了Audio-FLAN,一个大规模指令微调数据集,涵盖语音、音乐和声音领域的80种多样化任务,包含超过1亿个实例。Audio-FLAN为统一的音视频-语言模型奠定了基础,使其能够在零样本条件下无缝处理各类音频的理解(如转录、理解)与生成(如语音、音乐、声音)任务。该数据集已公开于HuggingFace和GitHub。

原文摘要 · Abstract (English)

Recent advancements in audio tokenization have significantly enhanced the integration of audio capabilities into large language models (LLMs). However, audio understanding and generation are often treated as distinct tasks, hindering the development of truly unified audio-language models. While instruction tuning has demonstrated remarkable success in improving generalization and zero-shot learning across text and vision, its application to audio remains largely unexplored. A major obstacle is the lack of comprehensive datasets that unify audio understanding and generation. To address this, we introduce Audio-FLAN, a large-scale instruction-tuning dataset covering 80 diverse tasks across speech, music, and sound domains, with over 100 million instances. Audio-FLAN lays the foundation for unified audio-language models that can seamlessly handle both understanding (e.g., transcription, comprehension) and generation (e.g., speech, music, sound) tasks across a wide range of audio domains in a zero-shot manner. The Audio-FLAN dataset is available on HuggingFace and GitHub.

音频理解指令微调多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。