探究扩展模态能否真正实现多模态统一,发现现有方法存在语言能力下降和泛化不足问题。
Is Extending Modality The Right Path Towards Omni-Modality?
- 通过微调通用语言模型扩展新模态,研究其对核心语言能力的影响。
- 独立训练的模态模型合并后仍无法实现跨模态泛化,表现受限。
- 相比逐次扩展,同时扩展多模态未显著提升知识共享与泛化能力。
全模态语言模型(OLMs)旨在整合并推理多种输入模态(如文本、图像、视频、音频),同时保持强大的语言能力。尽管近期取得进展,现有模型尤其是开源模型距离真正的全模态仍有差距,难以在特定模态对之外泛化,处理多模态输入时性能也较弱。本文研究主流的模态扩展方法——即对现成语言模型在目标领域和语言数据上进行微调——并探讨三个关键问题:(1) 模态扩展是否损害核心语言能力?(2) 模型合并能否有效集成各自微调的模态专用模型以实现全模态?(3) 全模态扩展相比顺序扩展是否带来更好的知识共享与泛化?通过大量实验分析这些权衡,为当前方法实现真正全模态提供了深入见解。
原文摘要 · Abstract (English)
Omni-modal language models (OLMs) aim to integrate and reason over diverse input modalities--such as text, images, video, and audio--while maintaining strong language capabilities. Despite recent advancements, existing models, especially open-source ones, remain far from true omni-modality, struggling to generalize beyond the specific modality pairs they are trained on or to achieve strong performance when processing multi-modal inputs. We study the effect of extending modality, the dominant technique for training multimodal models, where an off-the-shelf language model is fine-tuned on target-domain and language data. Specifically, we investigate three key questions: (1) Does modality extension compromise core language abilities? (2) Can model merging effectively integrate independently fine-tuned modality-specific models to achieve omni-modality? (3) Does omni-modality extension lead to better knowledge sharing and generalization compared to sequential extension? Through extensive experiments, we analyze these trade-offs and provide insights into the feasibility of achieving true omni-modality using current approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。