让大模型听懂声音,突破语音与环境音的通用理解难题
Cryfish: On deep audio analysis with Large Language Models
- 用Transformer连接器融合WavLM音频特征与Qwen2模型
- 在Dynamic SUPERB Phase-2上多任务表现超越公开模型
- 适合需要音频理解能力的研究者和开发者
近期基于文本的大语言模型(LLMs)的革命性进展,推动了将此类模型能力拓展至多模态感知与理解任务。听觉是亟需融入大语言模型的核心能力之一。然而,如何有效将听觉能力集成到大模型中,仍面临跨语音与声音的复杂任务泛化挑战。为此,我们提出Cryfish——一种具备听觉能力的大语言模型。该模型通过基于Transformer的连接器,将WavLM音频编码器特征融入Qwen2模型,并采用专用训练策略适配多种音频任务。我们在专为听觉能力模型设计的新版Dynamic SUPERB Phase-2多任务基准上评估该模型,深入分析并详细对比其与现有公开模型的表现。
原文摘要 · Abstract (English)
The recent revolutionary progress in text-based large language models (LLMs) has contributed to the growth of interest in extending capabilities of such models to multimodal perception and understanding tasks. Hearing is an essential capability that is highly desired to be integrated into LLMs. However, effective integrating listening capabilities into LLMs is a significant challenge lying in generalizing complex auditory tasks across speech and sounds. To address these issues, we introduce Cryfish, our version of auditory-capable LLM. The model integrates WavLM audio-encoder features into Qwen2 model using a transformer-based connector. Cryfish is adapted to various auditory tasks through a specialized training strategy. We evaluate the model on the new Dynamic SUPERB Phase-2 comprehensive multitask benchmark specifically designed for auditory-capable models. The paper presents an in-depth analysis and detailed comparison of Cryfish with the publicly available models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。