构建细粒度音频生成评估基准,精准检验文本到音频的语义一致性。
AudioScape-TTA: A Structured Soundscape Benchmark for Fine-Grained Text-to-Audio Evaluation

- 基于事件密度与结构复杂度设计分层评估框架
- 覆盖2258组音视频对,含25707条细粒度问答评测项
- 比传统相似性指标更贴近人类对音频语义的判断
文本到音频(TTA)生成近年来在从自然语言描述合成逼真音频方面取得显著进展。然而,判断生成音频是否忠实满足复杂文本指令仍具挑战性。现有基准主要依赖全局相似性度量,难以揭示细粒度语义错误。为此,我们提出 extbf{AudioScape-TTA},一个结构化且复杂度感知的细粒度 TTA 评估基准。该基准通过模态感知的语义结构表征真实声景,并以事件密度和结构复杂度刻画生成难度。基于这些标注,我们构建了基于评分标准的音频-文本对齐评估框架,通过细粒度语义标准验证事件实现、声学属性及语音内容。基准包含2,258组音视频对与25,707条二元问答评测项,支持可扩展、可解释的TTA系统分析。对13个代表性开源模型的实验揭示其在细粒度属性控制、语音内容保留及组合声景生成方面存在持续局限。人工验证表明,我们的评分框架与人类语义判断具有更强一致性,优于传统全局相似性度量。
原文摘要 · Abstract (English)
Text-to-audio (TTA) generation has recently achieved remarkable progress in synthesizing realistic audio from natural language descriptions. However, determining whether generated audio faithfully satisfies complex textual instructions remains challenging. Existing benchmarks mainly rely on global similarity metrics, providing limited insight into fine-grained semantic failures. To address this limitation, we introduce \textbf{AudioScape-TTA}, a structured and complexity-aware benchmark for fine-grained TTA evaluation. AudioScape-TTA represents realistic soundscapes through modality-aware semantic structures and characterizes generation complexity using event density and structural complexity. Based on these annotations, we propose a rubric-based audio-grounded evaluation framework that verifies event realization, acoustic attributes, and speech content through fine-grained semantic criteria. The benchmark contains 2,258 audio-text pairs with 25,707 binary QA rubrics, enabling scalable and interpretable analysis of TTA systems. Experiments on 13 representative open-source TTA models reveal persistent limitations in fine-grained attribute control, speech-content preservation, and compositional soundscape generation. Human validation further demonstrates that our rubric-based evaluation achieves stronger alignment with human semantic judgments than conventional global similarity metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。