arXiv:2412.11409cs.CVcs.AI2024-12AAAI被引 9

用图像和深度信息融合,让语音更真实地模拟环境回响。

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech

  • 结合RGB与深度图,多尺度捕捉局部与全局空间特征。
  • 通过生成的环境描述引导局部理解,提升语音还原度。
  • 适合做沉浸式语音合成、虚拟现实等场景的应用研究者。

视觉文本转语音(VTTS)旨在以环境图像为提示,合成具有混响效果的语音。该任务的核心挑战在于从图像中理解空间环境。现有方法多聚焦于从RGB图像中提取全局空间信息,却忽略了局部和深度信息的重要性。为此,本文提出一种新颖的多模态、多尺度空间环境理解方案——M2SE-VTTS,用于实现沉浸式VTTS。该方法同时利用空间图像的RGB与深度信息,以获取更全面的空间表征;并通过多尺度建模,同步捕捉局部与全局空间知识。具体而言,先将RGB与深度图像分割为图像块,并借助Gemini生成的环境描述引导局部空间理解;随后,通过局部感知的全局空间理解模块整合多模态、多尺度特征。实验表明,该模型在客观与主观评估中均优于现有先进基线。代码与音频样本已公开:https://github.com/AI-S2-Lab/M2SE-VTTS。

原文摘要 · Abstract (English)

Visual Text-to-Speech (VTTS) aims to take the environmental image as the prompt to synthesize the reverberant speech for the spoken content. The challenge of this task lies in understanding the spatial environment from the image. Many attempts have been made to extract global spatial visual information from the RGB space of an spatial image. However, local and depth image information are crucial for understanding the spatial environment, which previous works have ignored. To address the issues, we propose a novel multi-modal and multi-scale spatial environment understanding scheme to achieve immersive VTTS, termed M2SE-VTTS. The multi-modal aims to take both the RGB and Depth spaces of the spatial image to learn more comprehensive spatial information, and the multi-scale seeks to model the local and global spatial knowledge simultaneously. Specifically, we first split the RGB and Depth images into patches and adopt the Gemini-generated environment captions to guide the local spatial understanding. After that, the multi-modal and multi-scale features are integrated by the local-aware global spatial understanding. In this way, M2SE-VTTS effectively models the interactions between local and global spatial contexts in the multi-modal spatial environment. Objective and subjective evaluations suggest that our model outperforms the advanced baselines in environmental speech generation. The code and audio samples are available at: https://github.com/AI-S2-Lab/M2SE-VTTS.

语音合成多模态空间理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。