arXiv:2608.03742cs.SDcs.AI2026-08综述

AI可凭文本、图像等输入生成高质量音效,提升影视游戏制作效率。

AI-Based Sound Effect Generation: A Narrative Review of Generative Models Across Input Modalities

论文配图:AI-Based Sound Effect Generation: A Narrative Review of Generative Models Across Input Modalities
图 1 · 摘自论文原文
  • 基于多模态输入(文本/图像/音频)的AI音效生成模型
  • 部分模型达到当前最佳效果,音质高且语义对齐
  • 适合音效设计、互动媒体开发者参考

音效在数字应用中对动作、事件和环境提示至关重要,常需高度变化与情境适配。人工智能驱动的音频生成模型迅速发展,有望改变音效合成与应用方式。本文综述近五年30篇同行评审论文,分析不同输入模态(文本、视觉、音频、多模态)对生成音效质量、可控性与情境相关性的影响。结果显示,多个模型在任务中实现最先进的性能,生成高保真、语义一致且日益时序连贯的音效。然而,仍存在挑战:复杂多事件场景下时间同步不足,客观指标与人类感知存在差距,可控性与生成多样性之间存在权衡。整体表明,AI音效生成正朝着更自适应、可扩展、情境感知的方向演进,对未来的音效设计流程与互动媒体应用具有重要意义。

原文摘要 · Abstract (English)

Sound effects play a crucial role in conveying actions, events, and environmental cues across digital applications, often requiring a high degree of variation and contextual adaptability. Artificial intelligence (AI)-driven audio generative models are rapidly growing in popularity and have the potential to transform the way sound is synthesized and used across various applications. In response to this growing momentum, this chapter reviews and analyzes recent AI-based generative models for sound effect synthesis, with a focus on how different input modalities (text, visual, audio, and multimodal) affect the quality, controllability, and contextual relevance of the generated audio. It examines 30 peer-reviewed articles sourced from Google Scholar, IEEE Xplore, and the ACM Digital Library, exploring the evolution of AI generative models over the past five years. The results show that multiple models achieved state-of-the-art performance, producing high-fidelity, semantically aligned, and increasingly temporally coherent sound effects across tasks. However, despite these advances, the review identifies persistent challenges, including limitations in temporal synchronization for complex multi-event scenarios, gaps between objective metrics and human perception, and trade-offs between controllability and generative diversity. Overall, the chapter highlights that AI-driven sound effect generation is progressing toward more adaptive, scalable, and context-aware systems, offering significant implications for future sound design workflows and interactive media applications.

音效生成多模态AI音频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。