揭示统一多模态模型的伪融合现象及其信息熵差异根源
Pseudo-Unification: Entropy Probing Reveals Divergent Information Patterns in Unified Multimodal Models
- 用信息论框架分析输入编码与输出生成的联合过程
- 发现视觉与语言模态熵轨迹不同,文本生成与图像合成模式分裂
- 适合研究多模态模型内在机制的开发者与研究人员
统一多模态模型(UMMs)旨在结合大语言模型(LLMs)的推理能力与视觉模型的生成能力。然而实践中,这种协同作用仍难以实现:UMMs无法将类语言模型的推理能力迁移到图像生成中,且响应行为呈现显著差异。我们称此为伪融合。现有探针方法或缺乏模型内部洞察,或忽略提示-响应依赖关系。为此,我们提出一种信息论探针框架,联合分析UMMs如何编码输入并生成输出。应用于十种代表性UMMs,该框架揭示伪融合源于双重偏差:(i) 模态不对称编码,即视觉与语言遵循不同的熵演化路径;(ii) 模式分裂响应,即文本生成表现出高熵创造性,而图像合成则强制低熵保真度。仅当双侧实现统一(如通过上下文预测)的模型才展现出更真实的融合,即使参数更少也能实现更强的基于推理的文本到图像生成。本工作首次提供对融合性的模型内探查,表明真正的多模态协同需一致的信息流,而非仅共享参数。
原文摘要 · Abstract (English)
Unified multimodal models (UMMs) were designed to combine the reasoning ability of large language models (LLMs) with the generation capability of vision models. In practice, however, this synergy remains elusive: UMMs fail to transfer LLM-like reasoning to image synthesis and exhibit divergent response behaviors. We term this phenomenon pseudo-unification. Diagnosing its internal causes is important, but existing probing methods either lack model-internal insight or ignore prompt-response dependencies. To address these limitations, we propose an information-theoretic probing framework that jointly analyzes how UMMs encode inputs and generate outputs. Applied to ten representative UMMs, our framework reveals that pseudo-unification stems from a dual divergence: (i) Modality-Asymmetric Encoding, where vision and language follow different entropy trajectories, and (ii) Pattern-Split Response, where text generation exhibits high-entropy creativity while image synthesis enforces low-entropy fidelity. Only models that unify both sides (e.g., via contextual prediction) achieve more genuine unification, enabling stronger reasoning-based text-to-image generation even with fewer parameters. Our work provides the first model-internal probing of unification, demonstrating that real multimodal synergy requires consistency in information flow, not just shared parameters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。