用自监督语音特征与大模型结合,实现语音转文本的高效翻译。
SparQLe: Speech Queries to Text Translation Through LLMs
- 通过模态适配器对齐语音特征与指令微调大模型。
- 在英语语音数据上验证,能有效保留语音语义内容。
- 适合需要语音理解与文本生成融合的应用场景。
随着大语言模型(LLMs)的兴起,将语音表示与之结合以实现更无缝的多模态处理和语音理解成为研究热点。本文提出一种新方法,将自监督语音表示与指令微调的大型语言模型结合,用于语音到文本的翻译。该方法利用模态适配器,基于英语语音数据将提取的语音特征与指令微调的LLM对齐。实验表明,该方法能有效保留输入语音的语义内容,为自监督语音模型与指令微调大模型之间构建了有效桥梁,展现出在多种语音理解应用中的潜力。
原文摘要 · Abstract (English)
With the growing influence of Large Language Models (LLMs), there is increasing interest in integrating speech representations with them to enable more seamless multi-modal processing and speech understanding. This study introduces a novel approach that combines self-supervised speech representations with instruction-tuned LLMs for speech-to-text translation. The proposed approach leverages a modality adapter to align extracted speech features with instruction-tuned LLMs using English speech data. Our experiments demonstrate that this method effectively preserves the semantic content of the input speech and serves as an effective bridge between self-supervised speech models and instruction-tuned LLMs, offering a promising approach for various speech understanding applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。