让语音模型真正把想法表达出来,实现自然情感的语音生成。
Bridging What the Model Thinks and How It Speaks: Expressive Speech Generation via Self-Aware Intent-Realization Alignment

- 用自感知机制从模型内部推断表达意图,无需外部标注。
- 在生成过程中动态调整语音,使语气与意图一致,提升表现力。
- 小模型也能达到顶尖水平,适合追求真实感的语音应用。
语音语言模型(SLMs)具备强大的语义理解能力,但常无法将其转化为富有表现力的语音输出,导致语调平直、情绪错位。我们识别出这一问题为语义理解与语音实现之间的差距。现有方法多依赖外部标注的代理信号(如情绪标签或风格提示),需人工标注且难以捕捉对话中动态变化的表达意图。为此,我们提出SASLM(自感知语音语言模型),一种无代理框架,通过自感知意图-实现对齐来弥合“模型所想”与“语音所言”的鸿沟:(1) 意图感知桥接利用变分信息瓶颈(VIB)从模型自身不断演化的语义生成状态中自蒸馏表达意图,引导语音生成而无需外部表达监督;(2) 实现感知对齐通过自我奖励优化,反思性地将生成的语音与预期表达对齐,逐步提升生成过程中的意图-实现一致性。尽管仅使用30亿参数和800小时表达性语音数据,SASLM在开放源码系统中于EchoMind评测上达到最先进性能,超越参数量超十倍的模型,逼近商业系统水平。
原文摘要 · Abstract (English)
Speech Language Models (SLMs) exhibit strong semantic understanding, yet often fail to translate this capacity into expressive acoustic realization, producing speech with flattened prosody and misaligned emotion. We identify this mismatch as the semantic understanding-acoustic realization gap. Existing approaches typically rely on externally specified proxies, such as emotion labels or style prompts, which require annotations and struggle to capture dynamically evolving expressive intent throughout dialogue. To overcome these limitations, we propose SASLM (Self-Aware Speech Language Model), a proxy-free framework that bridges what the model thinks and how it speaks through self-aware intent-realization alignment: (1) Intent-Aware Bridging self-distills expressive intent from the model's own evolving semantic generation states via a Variational Information Bottleneck (VIB), thereby guiding expressive speech realization without external expressive supervision; while (2) Realization-Aware Alignment reflectively aligns generated acoustics with intended expression through self-reward optimization, progressively improving intent-realization consistency during speech generation. Despite using only 3B parameters and 800 hours of expressive speech data, SASLM achieves state-of-the-art performance on EchoMind among open-source systems, surpassing models over 10 times larger and approaching commercial systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。