arXiv:2606.23041cs.CV2026-06被引 1

让视觉理解与生成统一,实现无外部依赖的高质量图像重建。

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models

论文配图:SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models
图 1 · 摘自论文原文
  • 用双流令牌化器分离语义与像素特征,统一到紧凑潜在空间。
  • 自对齐生成机制使扩散模型无需外部教师即可精准还原细节。
  • 动态令牌路由支持灵活交互,适合多模态生成与理解任务。

多模态大语言模型在视觉理解上表现卓越,但在视觉生成方面受限于语义感知与像素级重建之间的根本差异。为解决这一问题,我们提出一种新型统一多模态框架——SPAR(语义-像素自对齐与自适应路由)。首先,设计非对称双流统一令牌化器:轻量语义流保留判别性特征,而基于Transformer的像素流恢复细粒度视觉细节,统一至紧凑潜在空间。其次,提出自对齐生成范式,利用优化后的令牌化器作为内部对齐教师,使扩散模型无需外部指导即可实现高保真重建。此外,引入动态令牌路由,使每个令牌根据自身语义需求自适应聚合多层特征。大量实验表明,SPAR在统一架构中达到当前最佳性能,兼具优异的生成与重建质量,同时保持基础视觉理解能力。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have achieved remarkable success in visual understanding but remain constrained in visual generation due to the fundamental feature discrepancy between semantic perception and pixel-level reconstruction. Bridging this gap requires overcoming two core challenges: endowing semantic encoders with high-fidelity reconstruction capabilities, and effectively aligning generative models with semantic spaces without relying on external teachers. To this end, we propose a novel unified multimodal framework featuring \textbf{S}emantic-\textbf{P}ixel self-alignment and \textbf{A}daptive \textbf{R}outing (\textbf{SPAR}). First, to reconcile semantic perception with pixel-level reconstruction, we introduce an asymmetric dual-stream unified tokenizer. A lightweight semantic stream anchors discriminative features, while a Transformer-augmented pixel stream recovers fine-grained visual details into a unified compact latent space. Second, to eliminate external dependencies, we propose a self-aligned generation paradigm that natively leverages this optimized tokenizer as an internal alignment teacher for the diffusion model. Furthermore, to facilitate flexible multimodal interaction within this unified space, we introduce Dynamic Token Routing, which enables each token to adaptively aggregate multi-layer MLLM features based on its distinct semantic demands. Extensive experiments demonstrate that SPAR establishes the state-of-the-art for unified architectures, achieving exceptional generation and reconstruction quality while preserving foundational visual understanding capabilities.

多模态图像生成统一模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。