用嵌入空间信号融合多模态模型,提升跨模态推理与单模态保留能力
ES-Merging: Biological MLLM Merging via Embedding Space Signals
- 基于嵌入空间信号估算融合系数,替代传统参数空间方法
- 在跨模态推理和单模态知识保留上均优于现有方法
- 适合需要多模态融合的生物科学发现场景
生物多模态大语言模型(MLLM)已成为科学发现的强大基础模型。然而,现有模型仅针对单一模态,难以解决本质上的跨模态科学问题。尽管模型融合是高效整合不同模态为统一MLLM的方法,但现有方法依赖输入无关的参数空间启发式策略,无法准确捕捉模态专长。为此,我们提出基于嵌入信号的MLLM融合框架ES-Merging,将融合范式从参数信号转向嵌入信号。ES-Merging利用嵌入空间中的粗粒度和细粒度信号,分别估计层间与元素级融合系数,并联合优化实现互补融合。大量实验表明,ES-Merging不仅在跨模态推理上表现更优,还在单模态知识保留方面显著领先,证明嵌入空间信号为MLLM融合提供了原则性且有效的基础。
原文摘要 · Abstract (English)
Biological multimodal large language models (MLLMs) have emerged as powerful foundation models for scientific discovery. However, existing models are specialized to a single modality, limiting their ability to solve inherently cross-modal scientific problems. While model merging is an efficient method to combine the different modalities into a unified MLLM, existing methods rely on input-agnostic parameter space heuristics that fail to faithfully capture modality specialization. To overcome this limitation, we propose the Embedding-Signal-based MLLM Merging (ES-Merging), a framework that estimates merging coefficients from embedding space signals, moving the merging paradigm from the parameter signals to the embedding signals. ES-Merging exploits coarse-grained and fine-grained signals from embedding space to estimate the layer-wise and element-wise merging coefficients, respectively, which are jointly combined for complementary coefficient estimation. Through extensive experiments, we demonstrate that ES-Merging outperforms existing merging methods not only on the cross-modal reasoning but also on the single-modal knowledge preservation, establishing that embedding space signals provide a principled and effective foundation for MLLM merging.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。