arXiv:2511.03423eess.AScs.CV2025-11被引 2

用语音直接生成有表现力的图像,保留语气与情感细节。

Seeing What You Say: Expressive Image Generation from Speech

  • 直接处理语音令牌,融合语言与语调信息生成图像。
  • 在三个数据集上验证有效,但情绪一致性仍存挑战。
  • 适合语音交互、创意生成等需情感表达的场景。

本文提出 VoxStudio,首个统一且端到端的语音转图像模型,可直接从口语描述生成富有表现力的图像,同时对齐语言与副语言信息。核心是语音信息瓶颈(SIB)模块,将原始语音压缩为紧凑的语义令牌,保留语调与情感细微差别。通过直接操作这些令牌,VoxStudio 避免了依赖额外的语音转文本系统,后者常忽略语气、情感等文本之外的隐含信息。我们还发布 VoxEmoset,一个基于先进语音合成引擎构建的大规模情感语音-图像配对数据集,可低成本生成丰富表达的语音。在 SpokenCOCO、Flickr8kAudio 与 VoxEmoset 基准上的实验验证了方法可行性,并揭示关键挑战:情绪一致性与语言模糊性,为未来研究铺平道路。

原文摘要 · Abstract (English)

This paper proposes VoxStudio, the first unified and end-to-end speech-to-image model that generates expressive images directly from spoken descriptions by jointly aligning linguistic and paralinguistic information. At its core is a speech information bottleneck (SIB) module, which compresses raw speech into compact semantic tokens, preserving prosody and emotional nuance. By operating directly on these tokens, VoxStudio eliminates the need for an additional speech-to-text system, which often ignores the hidden details beyond text, e.g., tone or emotion. We also release VoxEmoset, a large-scale paired emotional speech-image dataset built via an advanced TTS engine to affordably generate richly expressive utterances. Comprehensive experiments on the SpokenCOCO, Flickr8kAudio, and VoxEmoset benchmarks demonstrate the feasibility of our method and highlight key challenges, including emotional consistency and linguistic ambiguity, paving the way for future research.

语音生成图像生成情感表达端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。