arXiv:2607.12706cs.SD2026-07被引 1

自动分离语音风格,支持任意风格替换而不失自然感。

AutoSIFT: Automatic Style Sifting for Controllable Speech Generation with Arbitrary Style Infilling

论文配图:AutoSIFT: Automatic Style Sifting for Controllable Speech Generation with Arbitrary Style Infilling
图 1 · 摘自论文原文
  • 将语音风格分为可描述类别与残留细节,实现精准拆解。
  • 仅替换指定风格类别,保留语调与说话人特征。
  • 适合影视配音、游戏配音等需精细控制的场景。

当前最先进的文本到语音(TTS)模型在自然度和表现力上表现优异,但对说话风格的细粒度、解耦控制仍具挑战。在电影配音、游戏配音和视频内容生成等专业场景中,用户常需修改特定风格属性(如情绪、年龄或性别),同时保持其他所有特征不变。现有风格可控的TTS方法通常依赖文本描述风格或语音参考风格迁移,难以同时控制显式语义属性并保留细微的、未被文本描述的语调细节。我们提出AutoSIFT,一种面向类别级风格编辑的可控语音生成框架。AutoSIFT将说话风格分解为可由文本描述的类别与捕捉非语言语调及说话人特性的未知残差风格。该框架包含一个通用风格解耦器,从参考语音中提取类别感知的风格原型;以及一个任意风格补全器,可选择性地从参考语音中补全未指定的风格类别。通过仅替换文本指定的风格类别而保留源自参考语音的残差风格,AutoSIFT实现了自然、富有表现力且高度可定制的语音生成。

原文摘要 · Abstract (English)

State-of-the-art text-to-speech (TTS) models achieve impressive naturalness and expressiveness, yet fine-grained, disentangled control over speaking styles remains challenging. In professional scenarios such as film dubbing, game voice acting, and video content generation, users often need to modify a specific style category, such as emotion, age, or gender, while preserving all others. Existing style-controllable TTS methods typically rely on either text-described styles or speech-reference style transfer, making it difficult to jointly control explicit semantic attributes and preserve subtle, text-undescribed prosodic details. We propose AutoSIFT, a controllable speech generation framework for category-level style editing. AutoSIFT decomposes speaking style into known text-describable categories and unknown residual styles that capture non-verbal prosody and speaker-specific nuances. It consists of a generalized Style Disentangler, which extracts category-aware style prototypes from reference speech, and an Arbitrary Style Infiller, which selectively infills unspecified style categories from the reference. By replacing only text-specified style categories while preserving residual speech-derived styles, AutoSIFT enables natural, expressive, and highly customizable speech generation.

语音生成风格控制可编辑性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。