arXiv:2510.27102cs.SDcs.AI2025-10中稿 · AAAI

量化文本生成音频的表达范围,揭示其多样性与一致性。

Expressive Range Characterization of Open Text-to-Audio Models

  • 基于固定提示分析音频模型输出,评估其表达范围。
  • 在ESC-50数据集提示下,测得音高、响度、音色等维度的可变性。
  • 为游戏音乐与多模态内容生成提供可量化的评估框架。

文本到音频模型是一类根据文本提示生成音频输出的生成模型。尽管程序化内容生成(PCG)领域普遍关注关卡生成及其功能属性(如可玩性),但能引发玩家情感共鸣的游戏往往融合多种创意与多模态内容(如音乐、音效、视觉、叙事基调),因此多模态模型已开始用于此类场景。然而,这类模型具体生成什么内容,其变异性与保真度如何仍不明确——音频作为生成系统的目标,其范畴极为广泛。在PCG社区中,表达范围分析(ERA)已被用作量化生成器输出空间的方法,尤其适用于关卡生成器。本文将ERA方法拓展至文本到音频模型,通过针对特定固定提示分析输出的表达范围,使分析更具可行性。实验采用来自环境声音分类(ESC-50)数据集的标准化提示,对生成音频在关键声学维度(如音高、响度、音色)上进行分析。更广泛而言,本文提出了一种基于ERA的探索性评估框架,可用于生成式音频模型的性能考察。

原文摘要 · Abstract (English)

Text-to-audio models are a type of generative model that produces audio output in response to a given textual prompt. Although level generators and the properties of the functional content that they create (e.g., playability) dominate most discourse in procedurally generated content (PCG), games that emotionally resonate with players tend to weave together a range of creative and multimodal content (e.g., music, sounds, visuals, narrative tone), and multimodal models have begun seeing at least experimental use for this purpose. However, it remains unclear what exactly such models generate, and with what degree of variability and fidelity: audio is an extremely broad class of output for a generative system to target. Within the PCG community, expressive range analysis (ERA) has been used as a quantitative way to characterize generators' output space, especially for level generators. This paper adapts ERA to text-to-audio models, making the analysis tractable by looking at the expressive range of outputs for specific, fixed prompts. Experiments are conducted by prompting the models with several standardized prompts derived from the Environmental Sound Classification (ESC-50) dataset. The resulting audio is analyzed along key acoustic dimensions (e.g., pitch, loudness, and timbre). More broadly, this paper offers a framework for ERA-based exploratory evaluation of generative audio models.

音频生成表达范围多模态评估框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。