构建多控制混合音频生成评测基准,推动复杂声景生成研究
MMAG: A Multi-Control Mixed Audio Generation Benchmark
- 设计包含4000段人工验证音频的多维度标注数据集
- 发现现有模型在语义一致性和时间控制上存在明显性能短板
- 适合音频生成、语音合成与多模态研究者使用
近期音频生成系统已从单一模态合成发展到生成包含语音、音乐和音效的复杂声景。因此,评估这些模型需同时衡量语义保真度、说话人一致性与时间控制能力,但现有基准多聚焦于孤立领域或粗粒度描述。为此,我们提出多控制混合音频生成(MMAG)基准。MMAG包含约4,000段人工验证的音频片段,具有丰富标注,涵盖语音内容、说话人身份、音乐属性、声音事件及时间关系,并设有专门用于语音克隆和时间戳条件生成的子集。我们进一步提出系统化评估协议,量化声学保真度、语音质量、语义对齐与时间准确性。对代表性智能体编排器、统一音视频生成模型和原生混合音频生成器的基准测试显示,各模型在多项能力间存在显著性能权衡,无一模型表现全面均衡。结果揭示了可控混合音频生成仍面临重大挑战,并确立了MMAG作为未来研究的综合性基准。
原文摘要 · Abstract (English)
Recent audio generation systems have progressed from single-modality synthesis to generating complex acoustic scenes containing speech, music, and sound effects. Therefore, evaluating these models requires assessing multiple interacting capabilities, including semantic fidelity, speaker consistency, and temporal control, yet existing benchmarks focus on isolated domains or coarse-grained descriptions. To address this gap, we introduce the Multi-control Mixed Audio Generation (MMAG) benchmark. MMAG contains approximately 4,000 manually verified audio clips with rich annotations covering speech content, speaker identity, music attributes, sound events, and temporal relationships, together with dedicated subsets for voice cloning and timestamp-conditioned generation. We further propose a systematic evaluation protocol that measures acoustic fidelity, speech quality, semantic alignment, and temporal accuracy. Benchmarking representative agentic orchestrators, unified audio-visual generation models, and native mixed-audio generators reveals substantial performance trade-offs across these capabilities, with no existing model performing consistently well. Our results highlight the remaining challenges of controllable mixed audio generation and establish MMAG as a comprehensive benchmark for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。