arXiv:2504.07153cs.SDcs.AR2025-04被引 1

AI可跨模态生成沉浸式音景,让声音与文本、动画自由转换。

Artificial intelligence in creating, representing or expressing an immersive soundscape

  • 利用AI分析多模态数据,实现文本、声音、动画间的自动转换。
  • 能预测和生成音景数据,提升虚拟现实中的听觉沉浸感。
  • 适合研究人机交互、数字音频创作的学者与开发者。

在当今技术驱动的世界中,人工智能与虚拟现实的进步显著。这些发展推动了二者在音景领域的交叉研究。不仅引发对如何用新技术设计与创造音景的思考,也深入探讨其对人类听觉环境感知、理解与表达的影响。本文旨在回顾并讨论人工智能在该领域的最新应用,探索如何利用其识别复杂数据模式的能力,构建虚拟现实沉浸式音景。重点在于跨模态转换(如文本→声音、声音→动画)及多域数据预测与生成。研究关注人工智能在音景数据预测、检测与理解方面的能力,最终目标是弥合声音与其他人类可读数据之间的鸿沟。

原文摘要 · Abstract (English)

In today's tech-driven world, significant advancements in artificial intelligence and virtual reality have emerged. These developments drive research into exploring their intersection in the realm of soundscape. Not only do these technologies raise questions about how they will revolutionize the way we design and create soundscapes, but they also draw significant inquiries into their impact on human perception, understanding, and expression of auditory environments. This paper aims to review and discuss the latest applications of artificial intelligence in this domain. It explores how artificial intelligence can be utilized to create a virtual reality immersive soundscape, exploiting its ability to recognize complex patterns in various forms of data. This includes translating between different modalities such as text, sounds, and animations as well as predicting and generating data across these domains. It addresses questions surrounding artificial intelligence's capacity to predict, detect, and comprehend soundscape data, ultimately aiming to bridge the gap between sound and other forms of human-readable data. 1.

AI音景虚拟现实跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。