通过语义初始化与规划损失,提升语音语言模型的声学一致性。
Optimizing Speech Language Models for Acoustic Consistency
- 用自监督特征初始化语音令牌,结合轻量对齐损失训练。
- 0.7B和1.0B语音模型在多种条件下一致性最优,优于更大模型。
- 适合关注语音生成稳定性与语义-声学对齐的研究者。
我们研究了融合语义初始化与规划损失的语音语言模型,以实现鲁棒且一致的生成效果。方法包括使用自监督特征初始化语音令牌,施加轻量级对齐损失,并通过剪枝与辅助目标提升鲁棒性与内容规划能力。训练了三个模型:一个0.7B语音单模模型、一个1.0B语音单模模型,以及一个1.0B交错模型(同时包含文本与语音)。声学实验表明,语音单模模型在说话人、性别、情感、房间及背景因素下均表现出最高一致性,优于更大规模系统。交错模型提升了词汇与句法探针表现及语义-声学对齐,但降低了一致性。线性探针显示,初始化使模型偏向内容结构,但牺牲了韵律细节。结果表明,仅通过语言模型设计与训练策略即可调控声学稳定性与语义基础之间的平衡,无需修改分词器或运行时架构。演示与模型权重已开放供探索。
原文摘要 · Abstract (English)
We study speech language models that incorporate semantic initialization and planning losses to achieve robust and consistent generation. Our approach initializes speech tokens with self-supervised features, applies a light alignment loss, and trains with thinning and auxiliary objectives that target robustness and content planning. We train three models: a 0.7B speech-only model, a 1.0B speech-only model, and a 1.0B interleaved model with both text and speech. Acoustic studies show that the speech-only models achieve the highest consistency across speaker, gender, sentiment, room, and background factors, surpassing larger systems. Interleaving improves lexical and syntactic probes and semantic--acoustic alignment but reduces consistency. Linear probes show that our initialization biases the model toward content structure while trading off prosody detail. These results show that LM-side design and training mix control the balance between acoustic stability and semantic grounding without changes to the tokenizer or runtime architecture. A demo and model weights are available for exploration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。