arXiv:2605.30965eess.AScs.AI2026-05ACL

让语音自然融入环境声,提升真实感与可懂度。

ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment

论文配图:ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment
图 1 · 摘自论文原文
  • 用多模态扩散变换器建模语音与环境声的跨模态交互
  • 在客观指标和人工评测中优于现有方法,自然度与清晰度更高
  • 适合虚拟现实、智能助手等需要沉浸式语音的场景

近期文本引导的音频生成技术在音效、语音和音乐等领域取得显著进展。然而,由于语音与环境声在声学特征和时间动态上存在本质差异,联合生成仍具挑战。本文提出ImmersiveTTS,一种环境感知的文本转语音模型,通过显式建模跨模态交互,将语义对齐的语音隐空间与文本条件化的环境上下文融合。采用联合注意力机制实现多模态信息融合,并引入针对环境感知语音任务设计的领域特定表示对齐目标,利用语音与音频编码器的互补自监督表示增强语义一致性。实验结果表明,ImmersiveTTS在客观指标和人工听觉测试中均优于现有方法,在自然度、可懂性和音频保真度方面表现更优。

原文摘要 · Abstract (English)

Recent advancements in text-guided audio generation have yielded promising results in diverse domains, including sound effects, speech, and music. However, jointly generating speech with environmental audio remains challenging due to the inherent disparities in their acoustic patterns and temporal dynamics. We propose ImmersiveTTS, an environment-aware text-to-speech (TTS) model that generates natural speech seamlessly integrated within environmental contexts by explicitly modeling cross-modal interactions. Our model builds on a multimodal diffusion transformer and fuses transcript-aligned speech latent with text-conditioned environmental context via joint attention. To enhance semantic consistency, we introduce a domain-specific representation alignment objective tailored to environment-aware TTS, leveraging complementary self-supervised representations from speech and audio encoders. Experimental results show that ImmersiveTTS achieves higher naturalness, intelligibility, and audio fidelity than existing approaches across objective metrics and human listening tests.

语音生成多模态扩散模型环境感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。