用持续预训练让语音大模型兼顾听懂和生成,首次实现纯编码器令牌端到端语音翻译。
Balancing Speech Understanding and Generation Using Continual Pre-training for Codec-based Speech LLM
- 通过持续预训练将文本大模型适配为处理编码离散语音表示
- 在语音识别、语音合成、语音翻译等任务上均取得优异表现
- 首个仅用神经编码器令牌的单次通过端到端语音翻译系统
近年来,语音语言模型(LLMs)将文本大模型拓展至语音领域,但如何平衡语音理解与生成仍是挑战,尤其在基于编码器的表示下。本文提出一种持续预训练(CPT)框架,将文本大模型适配为处理编码离散化语音,缓解模态错位问题并保持语言推理能力。所提统一模型支持理解与生成,在自动语音识别(ASR)、文本转语音(TTS)、语音到文本翻译(S2T-Trans)及语音到语音翻译(S2S-Trans)任务上均表现优异。特别地,我们首次实现了仅使用神经编码器令牌的端到端单次通过式语音翻译系统,无需中间转录、翻译或语义令牌。CPT对跨模态对齐和任务泛化至关重要,是构建鲁棒统一语音大模型的强大工具。
原文摘要 · Abstract (English)
Recent advances in speech language models (LLMs) have extended textual LLMs to the speech domain, but balancing speech understanding and generation remains challenging, especially with codec-based representations. We propose a continual pre-training (CPT) framework that adapts a textual LLM to handle codec-discretized speech, mitigating modality mismatch and preserving linguistic reasoning. Our unified model supports both understanding and generation, achieving strong results across ASR, TTS, S2T-Trans, and S2S-Trans. Notably, we present the first end-to-end, single-pass S2S-Trans system using only neural codec tokens, without intermediate transcriptions, translations, or semantic tokens. CPT proves essential for cross-modal alignment and task generalization, making it a powerful tool for building robust, unified speech LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。