arXiv:2509.17164cs.SDeess.AS2025-09

用语音直接生成音频,比传统方法快76.9%且更准。

STAR: Speech-to-Audio Generation via Representation Learning

  • 通过语音直接提取声音事件语义,实现端到端音频生成。
  • 语音处理延迟降低76.9%,生成效果优于串联系统。
  • 适合语音交互、智能助手等需实时音频生成场景。

本文提出STAR,首个端到端的语音到音频生成框架,旨在提升效率并解决串联系统中的误差传播问题。与依赖文本或视觉的方法不同,STAR利用语音这一自然交互模态。通过表示学习实验,我们验证了从原始语音中可有效提取声音事件语义,涵盖听觉事件与场景线索。基于该语义表示,STAR引入桥接网络进行表征映射,并采用两阶段训练策略实现端到端合成。在语音处理延迟减少76.9%的前提下,生成性能显著优于串联系统。总体上,STAR确立了语音作为音频生成的直接交互信号,推动了表示学习与多模态合成的融合。生成样例见https://zeyuxie29.github.io/STAR。

原文摘要 · Abstract (English)

This work presents STAR, the first end-to-end speech-to-audio generation framework, designed to enhance efficiency and address error propagation inherent in cascaded systems. Unlike prior approaches relying on text or vision, STAR leverages speech as it constitutes a natural modality for interaction. As an initial step to validate the feasibility of the system, we demonstrate through representation learning experiments that spoken sound event semantics can be effectively extracted from raw speech, capturing both auditory events and scene cues. Leveraging the semantic representations, STAR incorporates a bridge network for representation mapping and a two-stage training strategy to achieve end-to-end synthesis. With a 76.9% reduction in speech processing latency, STAR demonstrates superior generation performance over the cascaded systems. Overall, STAR establishes speech as a direct interaction signal for audio generation, thereby bridging representation learning and multimodal synthesis. Generated samples are available at https://zeyuxie29.github.io/STAR.

语音生成端到端表示学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。