通过噪声查询令牌增强视觉语言与扩散模型的连接
WeMMU: Enhanced Bridging of Vision-Language Models and Diffusion Models via Noisy Query Tokens
- 用可学习的噪声查询令牌建立视觉语言与扩散模型间的分布表示空间
- 在多任务持续学习中显著缓解泛化崩溃,保持稳定性能
- 适合需要跨模态生成和持续学习的开发者与研究者
近期多模态大语言模型的发展凸显了如何高效连接预训练视觉-语言模型(VLM)与扩散模型的挑战。现有使用固定数量可学习查询令牌的方法虽具计算效率,但存在任务泛化崩溃问题,难以适应与预训练任务差异较大的新任务。为此,我们提出噪声查询令牌,通过端到端优化学习VLM与扩散模型之间的分布式表示空间,增强持续学习能力。此外,引入一个带线性投影的变分自编码器分支,以恢复图像的细粒度细节。实验结果表明,该方法有效缓解泛化崩溃,在多样化任务上实现稳定的持续学习。
原文摘要 · Abstract (English)
Recent progress in multimodal large language models (MLLMs) has highlighted the challenge of efficiently bridging pre-trained Vision-Language Models (VLMs) with Diffusion Models. While methods using a fixed number of learnable query tokens offer computational efficiency, they suffer from task generalization collapse, failing to adapt to new tasks that are distant from their pre-training tasks. To overcome this, we propose Noisy Query Tokens, which learn a distributed representation space between the VLM and Diffusion Model via end-to-end optimization, enhancing continual learning. Additionally, we introduce a VAE branch with linear projection to recover fine-grained image details. Experimental results confirm our approach mitigates generalization collapse and enables stable continual learning across diverse tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。