arXiv:2511.17429cs.SDeess.AS2025-11

探索文字生成音乐的语义与符号互动,揭示AI如何重塑听觉认知。

Semantic and Semiotic Interplays in Text-to-Audio AI: Exploring Cognitive Dynamics and Musical Interactions

  • 从语言提示到声音输出,分析AI中的意义构建机制
  • 发现模型在音乐表征中实现稳定与颠覆的双重作用
  • 适合对音乐认知、AI创作感兴趣的学者与创作者

本文研究人工智能领域新兴的文字转音频范式,探讨其对音乐创作、解读与认知的变革性影响。通过结构主义与后结构主义视角,结合认知理论中的图式动态与元认知概念,分析当自然语言描述转化为复杂声学对象时,文本-音频模态间复杂的语义与符号互动过程。研究揭示了AI辅助音乐活动中的认知动态,包括图式同化与顺应、元认知反思及建构性感知。论文指出,文字转音频AI模型作为准音乐符号物,在稳定与解构传统形式之间游走,催生新的聆听方式与审美反思。以Udio为典型案例,研究显示该过程不仅生成新颖音乐表达,还促使听众发展批判性与结构性意识的聆听能力,深化对音乐结构、符号细微差别及社会文化背景的理解。最后提出,此类模型可成为认识论工具与准对象,推动音乐互动的根本转变,促进用户对音乐认知与文化基础的更深入把握。

原文摘要 · Abstract (English)

This paper investigates the emerging text-to-audio paradigm in artificial intelligence (AI), examining its transformative implications for musical creation, interpretation, and cognition. I explore the complex semantic and semiotic interplays that occur when descriptive natural language prompts are translated into nuanced sound objects across the text-to-audio modality. Drawing from structuralist and post-structuralist perspectives, as well as cognitive theories of schema dynamics and metacognition, the paper explores how these AI systems reconfigure musical signification processes and navigate established cognitive frameworks. The research analyzes some of the cognitive dynamics at play in AI-mediated musicking, including processes of schema assimilation and accommodation, metacognitive reflection, and constructive perception. The paper argues that text-to-audio AI models function as quasi-objects of musical signification, simultaneously stabilizing and destabilizing conventional forms while fostering new modes of listening and aesthetic reflexivity.Using Udio as a primary case study, this study explores how these models navigate the liminal spaces between linguistic prompts and sonic outputs. This process not only generates novel musical expressions but also prompts listeners to engage in forms of critical and "structurally-aware listening.", encouraging a deeper understanding of music's structures, semiotic nuances, and the socio-cultural contexts that shape our musical cognition. The paper concludes by reflecting on the potential of text-to-audio AI models to serve as epistemic tools and quasi-objects, facilitating a significant shift in musical interactions and inviting users to develop a more nuanced comprehension of the cognitive and cultural foundations of music.

文本生成音频音乐认知AI艺术

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。