用置信度和合成描述提升文本生成音频的质量。
CosyAudio: Improving Audio Generation with Confidence Scores and Synthetic Captions
- 通过置信度评估生成高质量合成描述,指导音频生成。
- 在多个数据集上生成更忠实于文本的音频,显著优于现有模型。
- 适合需要高保真音频生成与弱标注数据训练的研究者。
文本到音频(TTA)生成是AI生成内容(AIGC)中的新兴领域,旨在从自然语言描述生成音频。尽管关注度上升,但构建稳健的TTA模型仍面临挑战,主要源于高质量标注数据稀缺,以及大规模弱标注语料中普遍存在噪声或不准确的字幕。为此,我们提出CosyAudio框架,利用置信度评分和合成字幕来提升音频生成质量。该框架包含两个核心组件:AudioCapTeller和音频生成器。AudioCapTeller为音频生成合成字幕并提供置信度评分以评估其准确性;音频生成器则利用这些合成字幕和置信度实现质量感知的音频生成。此外,我们引入一种自进化训练策略,在高质量和弱标注数据集间迭代优化。初始在高质量数据上训练后,AudioCapTeller利用其评估能力对弱标注数据进行高质量筛选与强化学习,进一步提升性能。经过优化的AudioCapTeller可生成新字幕与置信度评分,用于音频生成器的训练。在开源数据集上的大量实验表明,CosyAudio在自动音频字幕任务中表现更优,生成的音频更忠实于输入文本,并在多种场景下展现出强泛化能力。
原文摘要 · Abstract (English)
Text-to-Audio (TTA) generation is an emerging area within AI-generated content (AIGC), where audio is created from natural language descriptions. Despite growing interest, developing robust TTA models remains challenging due to the scarcity of well-labeled datasets and the prevalence of noisy or inaccurate captions in large-scale, weakly labeled corpora. To address these challenges, we propose CosyAudio, a novel framework that utilizes confidence scores and synthetic captions to enhance the quality of audio generation. CosyAudio consists of two core components: AudioCapTeller and an audio generator. AudioCapTeller generates synthetic captions for audio and provides confidence scores to evaluate their accuracy. The audio generator uses these synthetic captions and confidence scores to enable quality-aware audio generation. Additionally, we introduce a self-evolving training strategy that iteratively optimizes CosyAudio across both well-labeled and weakly-labeled datasets. Initially trained with well-labeled data, AudioCapTeller leverages its assessment capabilities on weakly-labeled datasets for high-quality filtering and reinforcement learning, which further improves its performance. The well-trained AudioCapTeller refines corpora by generating new captions and confidence scores, serving for the audio generator training. Extensive experiments on open-source datasets demonstrate that CosyAudio outperforms existing models in automated audio captioning, generates more faithful audio, and exhibits strong generalization across diverse scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。