arXiv:2508.08961cs.SDeess.AS2025-08AAAI被引 4

用双语音标记统一语音理解和生成,提升模型性能。

DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language Models

  • 设计语义驱动的语音标记器,增强语音与文本模态对齐。
  • 提出双标记框架,同时支持语音理解与生成任务。
  • 适合需要统一语音处理能力的研究与应用。

通过引入有效语音标记扩展预训练文本大语言模型(LLM)的语音理解或生成能力,已成为语音领域的重要方向。然而,构建统一的语音理解与生成模型仍面临两大挑战:(1) 语音与文本标记之间存在巨大模态差距,将文本LLM扩展为统一语音LLM需依赖大规模配对数据微调;(2) 生成与理解任务偏好不同层次的信息,如生成需细致声学特征,理解则更关注高层语义。这种差异导致单一模型难以兼顾性能优化。本文提出两个关键洞察:首先,提出理解驱动的语音标记器(USTokenizer),利用文本LLM提取理解任务所需的高层语义信息,使USToken与文本具有更好模态共性,降低模态对齐难度;其次,提出DualSpeechLM双标记建模框架,在统一端到端架构中同时以USToken为输入、声学标记为输出,无缝集成语音理解与生成能力。此外,设计新型语义监督损失与链式条件(CoC)策略,稳定训练并提升生成性能。实验表明,该方法有效促进理解与生成任务间的互补关系,验证了在统一模型中相互增强的可行性。

原文摘要 · Abstract (English)

Extending pre-trained text Large Language Models (LLMs)'s speech understanding or generation abilities by introducing various effective speech tokens has attracted great attention in the speech community. However, building a unified speech understanding and generation model still faces the following challenges: (1) Due to the huge modality gap between speech and text tokens, extending text LLMs to unified speech LLMs relies on large-scale paired data for fine-tuning, and (2) Generation and understanding tasks prefer information at different levels, e.g., generation benefits from detailed acoustic features, while understanding favors high-level semantics. This divergence leads to difficult performance optimization in one unified model. To solve these challenges, in this paper, we present two key insights in speech tokenization and speech language modeling. Specifically, we first propose an Understanding-driven Speech Tokenizer (USTokenizer), which extracts high-level semantic information essential for accomplishing understanding tasks using text LLMs. In this way, USToken enjoys better modality commonality with text, which reduces the difficulty of modality alignment in adapting text LLMs to speech LLMs. Secondly, we present DualSpeechLM, a dual-token modeling framework that concurrently models USToken as input and acoustic token as output within a unified, end-to-end framework, seamlessly integrating speech understanding and generation capabilities. Furthermore, we propose a novel semantic supervision loss and a Chain-of-Condition (CoC) strategy to stabilize model training and enhance speech generation performance. Experimental results demonstrate that our proposed approach effectively fosters a complementary relationship between understanding and generation tasks, highlighting the promising strategy of mutually enhancing both tasks in one unified model.

语音生成大模型双标记统一建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。