arXiv:2607.26411cs.CV2026-07

通过跨分支操控,揭示统一多模态模型的语义空间是否真正统一。

Do Unified Multimodal Models Think in One Space? A Lens Through Cross-Branch Steering

论文配图:Do Unified Multimodal Models Think in One Space? A Lens Through Cross-Branch Steering
图 1 · 摘自论文原文
  • 用跨分支语义操控技术,将理解分支的语义方向迁移到生成分支。
  • 理解分支的向量可控制图像生成并提升语义准确度,反向则效果差。
  • 发现理解侧更捕捉对象级语义,生成侧主要编码外观细节。

统一多模态模型(UMMs)旨在以单一架构整合理解和生成能力,但其是否共享统一且可迁移的语义空间仍不明确。由于理解与生成分支分别处理文本标记和视觉潜变量,且训练目标不同,直接比较极为困难。为此,我们提出跨分支语义操控框架,从一个分支提取语义方向并施加于另一分支。实验表明,理解分支学习的操控向量可有效迁移至生成分支,实现可控图像生成并提升语义忠实度;而反向迁移始终效果有限。分析显示,这种不对称性源于表征失配:理解分支的向量捕捉可迁移的对象级语义,而生成分支的向量主要编码低层外观特征。结果表明,架构统一不等于语义对齐,并确立跨分支操控为探测多模态表征的有效工具。

原文摘要 · Abstract (English)

Unified multimodal models (UMMs) aim to integrate understanding and generation within a single architecture, yet it remains unclear whether these capabilities share a unified and transferable semantic space. This question is fundamentally challenging, as the two branches operate over heterogeneous representations (text tokens vs.\ visual latents) and distinct training objectives, making direct comparison difficult. To address this, we introduce \emph{cross-branch semantic steering}, an intervention-based framework that extracts semantic directions from one branch and applies them to the other. We show that steering vectors learned from the understanding branch can transfer to generation, enabling controllable image synthesis and improved semantic faithfulness. In contrast, the reverse direction consistently shows limited effectiveness. Our analysis suggests that this asymmetry may be related to a practical representational mismatch: understanding-derived vectors capture transferable, object-centric semantics, while generation-derived vectors primarily encode low-level appearance features. Our results reveal that architectural unification does not guarantee semantic alignment, and establish cross-branch steering as a practical tool for probing multimodal representations.

多模态语义空间可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。