arXiv:2410.14101cs.SDcs.AI2024-10中稿 · ICASSP'2025被引 10

融合深度、位置与语义信息,让语音更沉浸真实。

Multi-Source Spatial Knowledge Understanding for Immersive Visual Text-to-Speech

  • 用RGB+深度+位置+语义多源信息建模环境空间
  • 动态融合各源贡献,生成更具沉浸感的语音
  • 适合做虚拟现实语音合成的研究者参考

视觉文本转语音(VTTS)旨在以环境图像为提示,合成具有混响效果的口语语音。现有方法主要依赖RGB图像进行全局环境建模,忽略了深度图、说话人位置和环境语义等多源空间知识的潜力。为此,我们提出一种新的多源空间知识理解方案——MS2KU-VTTS。首先以RGB图像为主导源,将深度图、基于目标检测的说话人位置信息以及Gemini生成的语义描述作为补充源;随后设计串行交互机制,有效整合主导与补充源信息,并根据各源贡献动态融合。该增强后的多源空间知识引导语音生成模型,显著提升语音沉浸感。实验表明,MS²KU-VTTS在生成沉浸式语音方面优于现有基线模型。演示与代码已开源:https://github.com/AI-S2-Lab/MS2KU-VTTS。

原文摘要 · Abstract (English)

Visual Text-to-Speech (VTTS) aims to take the environmental image as the prompt to synthesize reverberant speech for the spoken content. Previous works focus on the RGB modality for global environmental modeling, overlooking the potential of multi-source spatial knowledge like depth, speaker position, and environmental semantics. To address these issues, we propose a novel multi-source spatial knowledge understanding scheme for immersive VTTS, termed MS2KU-VTTS. Specifically, we first prioritize RGB image as the dominant source and consider depth image, speaker position knowledge from object detection, and Gemini-generated semantic captions as supplementary sources. Afterwards, we propose a serial interaction mechanism to effectively integrate both dominant and supplementary sources. The resulting multi-source knowledge is dynamically integrated based on the respective contributions of each source.This enriched interaction and integration of multi-source spatial knowledge guides the speech generation model, enhancing the immersive speech experience. Experimental results demonstrate that the MS$^2$KU-VTTS surpasses existing baselines in generating immersive speech. Demos and code are available at: https://github.com/AI-S2-Lab/MS2KU-VTTS.

语音合成多模态沉浸式视觉生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。