arXiv:2607.11364cs.LGcs.AI2026-07

让长篇故事生成连贯影视级音效,自动对齐时间与情绪。

BackgroundMellow: A Multi-Modal Cohesive Framework for Narrative-Driven Rich Cinematic Soundscape Generation

论文配图:BackgroundMellow: A Multi-Modal Cohesive Framework for Narrative-Driven Rich Cinematic Soundscape Generation
图 1 · 摘自论文原文
  • 分层提取文本音频线索,用专用模型生成各类音效
  • 通过近邻检索验证,音效时间对齐度提升显著
  • 适合影视音效自动化、多模态内容生成研究者

为长篇文本叙事生成沉浸式、同步且具有电影感的音频仍是多模态AI的重大挑战。现有文本到音频(TTA)框架虽能合成孤立音效,但在叙事连贯性、时间对齐和情感深度上表现不足。本文提出BackgroundMellow,将故事转音频视为精准编排与信号处理问题。该框架采用无真实标注的主-专精代理架构:将文本分解为多层次音频线索,由对应专精模型生成各类声音,并叠加形成统一对齐的音景。基于Tango2潜在扩散模型实现环境音合成,结合从专业配乐中挖掘的新型电影背景音乐检索器(Cinematic BGM Retriever)。通过基于NLP的模块预测音效起始时间、持续时长和相对音量等参数,实现自动混音。在精选的YouTube电影预告片数据集上,采用最近邻检索方法评估,验证了框架在时间同步性、覆盖范围与频谱丰富度方面的有效性。

原文摘要 · Abstract (English)

Generating immersive, synchronized and cinematic audio for long-form textual narratives remains a significant challenge in multi-modal AI. While current Text-to-Audio (TTA) frameworks successfully synthesize isolated sound effects, they struggle with narrative cohesion, temporal alignment, and cinematic emotional depth. We present BackgroundMellow, a framework that treats story-to-audio generation as a precise orchestration and signal processing problem. This framework is enabled without ground-truth through a master-specialist agent architecture that decomposes text into precise and multi-layered audio cues, generates each category of sounds with suitable specialist model, and superimposes the soundscapes to create a unified and aligned audio segment. Our pipeline is built over Tango2 latent diffusion model for environmental synthesis alongside a novel Cinematic BGM Retriever mined from professional soundtracks. To automate the sound mixing process, we use an NLP based module that predicts precise audio parameters, like start time, duration, and relative loudness, based on the narrative timeline. We further empirically evaluate and show the efficacy of the proposed framework leveraging nearest-neighbor retrieval against a curated dataset of YouTube cinematic trailers to measure temporal synchronization, coverage, and spectral richness.

音效生成多模态文本转音频影视音景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。