让专家模型更懂任务意图,减少图像生成与编辑的冲突
TAG-MoE: Task-Aware Gating for Unified Generative Mixture-of-Experts
- 用语义标注和预测对齐正则化,让门控网络理解任务整体意图
- 在统一模型中同时提升生成质量和编辑精度,优于密集模型
- 适合需要多任务协同的图像生成与编辑场景
统一的图像生成与编辑模型在密集扩散变换器架构中面临严重的任务干扰问题,共享参数空间需在冲突目标间妥协(如局部编辑与主体驱动生成)。尽管稀疏的专家混合(MoE)范式是潜在解决方案,但其门控网络仍为任务无关型,仅基于局部特征决策,无法感知全局任务意图。这种任务无关性阻碍了有意义的专家专业化,未能解决根本的任务干扰。本文提出新框架,将语义意图注入MoE路由。引入分层任务语义标注方案,构建结构化任务描述符(如范围、类型、保留)。设计预测对齐正则化,使内部路由决策与任务高层语义对齐。该正则化使门控网络从任务无关执行者转变为调度中心。模型有效缓解任务干扰,在保真度与质量上超越密集基线;分析显示专家自然形成清晰且语义相关的专长。
原文摘要 · Abstract (English)
Unified image generation and editing models suffer from severe task interference in dense diffusion transformers architectures, where a shared parameter space must compromise between conflicting objectives (e.g., local editing v.s. subject-driven generation). While the sparse Mixture-of-Experts (MoE) paradigm is a promising solution, its gating networks remain task-agnostic, operating based on local features, unaware of global task intent. This task-agnostic nature prevents meaningful specialization and fails to resolve the underlying task interference. In this paper, we propose a novel framework to inject semantic intent into MoE routing. We introduce a Hierarchical Task Semantic Annotation scheme to create structured task descriptors (e.g., scope, type, preservation). We then design Predictive Alignment Regularization to align internal routing decisions with the task's high-level semantics. This regularization evolves the gating network from a task-agnostic executor to a dispatch center. Our model effectively mitigates task interference, outperforming dense baselines in fidelity and quality, and our analysis shows that experts naturally develop clear and semantically correlated specializations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。