测试文本生成音乐时对微小语义变化的敏感性,发现大模型仍易出错。
Evaluating Semantic Fragility in Text-to-Audio Generation Systems Under Controlled Prompt Perturbations
- 设计75组语义一致但表达微调的提示词,系统评估模型鲁棒性。
- 大模型在语义相似度上达0.82,但音频和时间特征仍差异明显。
- 适合关注生成音频稳定性的研究人员或应用开发者。
文本到音频生成技术虽能将自然语言描述转化为多样音乐,但其在语义等价提示词变化下的鲁棒性尚未充分研究。微小语言变化可能导致生成音频显著差异,影响实际可靠性。本研究针对MusicGen-small、MusicGen-large和Stable Audio 2.5三类代表性模型,通过最小词汇替换(MLS)、强度转移(IS)和结构重述(SR)进行受控扰动评估。构建包含75个提示组的数据集,保持语义意图的同时引入局部语言变异。采用谱域、时域与语义相似性多维度对比生成结果,实现跨表征层级的鲁棒性分析。实验表明,大模型在语义一致性上表现更优,MusicGen-large在MLS下达到0.77、IS下达0.82的余弦相似度;然而音色与时间特征分析显示,所有模型仍存在持续偏差,即便嵌入相似度高。结果表明,脆弱性主要源于语义到声学的转化过程,而非多模态对齐。本研究提出可控评估框架,强调生成音频系统需多层级稳定性检验。
原文摘要 · Abstract (English)
Recent advances in text-to-audio generation enable models to translate natural-language descriptions into diverse musical output. However, the robustness of these systems under semantically equivalent prompt variations remains largely unexplored. Small linguistic changes may lead to substantial variation in generated audio, raising concerns about reliability in practical use. In this study, we evaluate the semantic fragility of text-to-audio systems under controlled prompt perturbations. We selected MusicGen-small, MusicGen-large, and Stable Audio 2.5 as representative models, and we evaluated them under Minimal Lexical Substitution (MLS), Intensity Shifts (IS), and Structural Rephrasing (SR). The proposed dataset contains 75 prompt groups designed to preserve semantic intent while introducing localized linguistic variation. Generated outputs are compared through complementary spectral, temporal, and semantic similarity measures, enabling robustness analysis across multiple representational levels. Experimental results show that larger models achieve improved semantic consistency, with MusicGen-large reaching cosine similarities of 0.77 under MLS and 0.82 under IS. However, acoustic and temporal analyses reveal persistent divergence across all models, even when embedding similarity remains high. These findings indicate that fragility arises primarily during semantic-to-acoustic realization rather than multi-modal embedding alignment. Our study introduces a controlled framework for evaluating robustness in text-to-audio generation and highlights the need for multi-level stability assessment in generative audio systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。