arXiv:2412.05951cs.SDcs.CV2024-12

用轻量适配器让视觉模型直接懂音频,无需大量音频预训练。

When Vision Models Meet Parameter Efficient Look-Aside Adapters Without Large-Scale Audio Pretraining

  • 设计跨时频维度的适配器,增强视觉模型对音频特征的捕捉。
  • 在多个音频任务上达到甚至超越专用音频模型性能。
  • 适合资源有限但需快速部署视觉模型处理音频的场景。

近期研究显示,预训练视觉模型可提升音频下游任务表现。为进一步优化,通常需通过大规模音频数据进行额外预训练以注入音频特异性知识,但这要求海量音频数据和精心设计的目标函数。本文提出绕过预训练阶段,直接通过专为高效音频理解设计的「旁路适配器」(LoAA)微调视觉模型。音频频谱数据具有时间与频率双重异构维度,我们优化适配器以促进跨维度的标记交互。实验表明,该方法使视觉模型在多种音频与语音任务中达到或超越预训练音频模型性能,为利用视觉模型开展音频应用提供了一种资源高效且有效的解决方案。

原文摘要 · Abstract (English)

Recent studies show that pretrained vision models can boost performance in audio downstream tasks. To enhance the performance further, an additional pretraining stage with large scale audio data is typically required to infuse audio specific knowledge into the vision model. However, such approaches require extensive audio data and a carefully designed objective function. In this work, we propose bypassing the pretraining stage by directly fine-tuning the vision model with our Look Aside Adapter (LoAA) designed for efficient audio understanding. Audio spectrum data is represented across two heterogeneous dimensions time and frequency and we refine adapters to facilitate interactions between tokens across these dimensions. Our experiments demonstrate that our adapters allow vision models to reach or surpass the performance of pretrained audio models in various audio and speech tasks, offering a resource efficient and effective solution for leveraging vision models in audio applications.

视觉模型音频理解轻量化适配跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。