PicoAudio2实现自然语言控制的高质量音频生成,支持自由文本与时间精准调控。
PicoAudio2: Temporal Controllable Text-to-Audio Generation with Natural Language Description
- 用标注工具构建真实音频的时间戳数据,结合仿真数据训练
- 在真实数据集上生成音频质量显著提升,时间控制更精准
- 适合需要自由文本输入和精确时间控制的音频生成任务
当前可控文本到音频(TTA)生成方法虽通过时间戳条件实现细粒度控制,但受限于音频质量和输入格式。现有模型多依赖合成数据,导致真实数据下音频质量差;部分模型仅支持封闭词汇的声音事件,无法处理开放自由文本。本文提出PicoAudio2框架,通过接地模型标注真实音频-文本数据集中的事件时间戳,构建具有强时间信息的真实数据,并与已有仿真数据结合训练。模型采用改进架构,融合时间矩阵的细粒度信息与自由文本的粗粒度输入。实验表明,PicoAudio2在时间可控性和音频质量上均表现更优。
原文摘要 · Abstract (English)
While recent work in controllable text-to-audio (TTA) generation has achieved fine-grained control through timestamp conditioning, its scope remains limited by audio quality and input format. These models often suffer from poor audio quality in real datasets due to sole reliance on synthetic data. Moreover, some models are constrained to a closed vocabulary of sound events, preventing them from controlling audio generation for open-ended, free-text queries. This paper introduces PicoAudio2, a framework that advances temporal-controllable TTA by mitigating these data and architectural limitations. Specifically, we use a grounding model to annotate event timestamps of real audio-text datasets to curate temporally-strong real data, in addition to simulation data from existing works. The model is trained on the combination of real and simulation data. Moreover, we propose an enhanced architecture that integrates the fine-grained information from a timestamp matrix with coarse-grained free-text input. Experiments show that PicoAudio2 exhibits superior performance in terms of temporal controllability and audio quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。