让冻结的LLM学会看图说话,还能生成图像。
X-Fusion: Introducing New Modality to Frozen Large Language Models

- 用双塔结构保留语言模型,只加视觉专用参数。
- 图文任务表现优于其他方法,小模型收敛更快。
- 理解数据提升生成质量,降噪数据增强整体性能。
我们提出X-Fusion框架,使预训练大型语言模型(LLMs)在保持语言能力的同时拓展多模态任务能力。X-Fusion采用双塔设计,配备模态特异性权重,在冻结LLM参数的前提下整合视觉信息,实现理解和生成。实验表明,X-Fusion在图像到文本和文本到图像任务上均持续优于其他架构。研究发现:引入以理解为导向的数据可提升生成质量,降低图像数据噪声能增强整体性能,特征对齐可加速小模型收敛,但对大模型影响甚微。这些发现为构建高效统一的多模态模型提供了重要启示。
原文摘要 · Abstract (English)
We propose X-Fusion, a framework that extends pretrained Large Language Models (LLMs) for multimodal tasks while preserving their language capabilities. X-Fusion employs a dual-tower design with modality-specific weights, keeping the LLM's parameters frozen while integrating vision-specific information for both understanding and generation. Our experiments demonstrate that X-Fusion consistently outperforms alternative architectures on both image-to-text and text-to-image tasks. We find that incorporating understanding-focused data improves generation quality, reducing image data noise enhances overall performance, and feature alignment accelerates convergence for smaller models but has minimal impact on larger ones. Our findings provide valuable insights into building efficient unified multimodal models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。