arXiv:2602.03510cs.CV2026-02被引 4

通过分层语义路由提升扩散模型的文本生成能力

Semantic Routing: Exploring Multi-Layer LLM Feature Weighting for Diffusion Transformers

  • 设计轻量级门控机制,动态融合多层LLM特征
  • 深度路由策略使图文对齐提升9.97分(GenAI-Bench)
  • 适合追求高质量可控生成的研究者与开发者

基于DiT的文生图模型越来越多地采用LLM作为文本编码器,但文本条件通常静态且仅依赖单层LLM,忽略了LLM层间显著的语义层次结构以及扩散过程随时间和网络深度变化的非平稳去噪特性。为更好匹配DiT生成的动态过程,我们提出一种统一的归一化凸融合框架,配备轻量级门控,实现时间、深度及联合维度的多层LLM隐藏状态系统性整合。实验表明,深度语义路由是更优的条件策略,在图文对齐和组合生成任务上表现优异(如在GenAI-Bench Counting任务上提升9.97分)。相反,纯时间融合反而降低视觉生成保真度,原因在于训练-推理轨迹不匹配:在无分类器引导下,名义时间步无法跟踪有效信噪比,导致推理时语义错时注入。总体结果表明,深度路由可作为强大有效的基线,并强调了轨迹感知信号对实现鲁棒时间依赖条件生成的关键作用。

原文摘要 · Abstract (English)

Recent DiT-based text-to-image models increasingly adopt LLMs as text encoders, yet text conditioning remains largely static and often utilizes only a single LLM layer, despite pronounced semantic hierarchy across LLM layers and non-stationary denoising dynamics over both diffusion time and network depth. To better match the dynamic process of DiT generation and thereby enhance the diffusion model's generative capability, we introduce a unified normalized convex fusion framework equipped with lightweight gates to systematically organize multi-layer LLM hidden states via time-wise, depth-wise, and joint fusion. Experiments establish Depth-wise Semantic Routing as the superior conditioning strategy, consistently improving text-image alignment and compositional generation (e.g., +9.97 on the GenAI-Bench Counting task). Conversely, we find that purely time-wise fusion can paradoxically degrade visual generation fidelity. We attribute this to a train-inference trajectory mismatch: under classifier-free guidance, nominal timesteps fail to track the effective SNR, causing semantically mistimed feature injection during inference. Overall, our results position depth-wise routing as a strong and effective baseline and highlight the critical need for trajectory-aware signals to enable robust time-dependent conditioning.

扩散模型语义路由文本生成LLM融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。