arXiv:2603.27666cs.CV2026-03

提出新框架,让轻量模型在手机上也能精准控图。

Gated Condition Injection without Multimodal Attention: Towards Controllable Linear-Attention Transformers

  • 用双路门控模块统一处理多种条件输入
  • 在多任务测试中超越现有线性注意力模型的生成质量
  • 适合移动端部署,兼顾隐私与效率

基于扩散模型的可控图像生成近期取得显著进展,但因其计算量大,通常依赖云端部署,引发用户数据隐私担忧。为实现安全高效的本地生成,本文探索基于线性注意力架构的可控扩散模型,该架构具备优异可扩展性与效率,适用于边缘设备。然而实验发现,现有可控生成框架如ControlNet和OminiControl在该类模型上或无法支持多种异构条件输入,或收敛速度慢。为此,本文提出一种专为SANA等线性注意力骨干网络设计的新型可控扩散框架。核心是一个统一的门控条件注入模块,采用双路径结构,有效融合空间对齐与非对齐条件输入。大量实验表明,该方法在多个任务与基准上达到线性注意力模型上的最优可控生成性能,显著提升保真度与可控性。

原文摘要 · Abstract (English)

Recent advances in diffusion-based controllable visual generation have led to remarkable improvements in image quality. However, these powerful models are typically deployed on cloud servers due to their large computational demands, raising serious concerns about user data privacy. To enable secure and efficient on-device generation, we explore in this paper controllable diffusion models built upon linear attention architectures, which offer superior scalability and efficiency, even on edge devices. Yet, our experiments reveal that existing controllable generation frameworks, such as ControlNet and OminiControl, either lack the flexibility to support multiple heterogeneous condition types or suffer from slow convergence on such linear-attention models. To address these limitations, we propose a novel controllable diffusion framework tailored for linear attention backbones like SANA. The core of our method lies in a unified gated conditioning module working in a dual-path pipeline, which effectively integrates multi-type conditional inputs, such as spatially aligned and non-aligned cues. Extensive experiments on multiple tasks and benchmarks demonstrate that our approach achieves state-of-the-art controllable generation performance based on linear-attention models, surpassing existing methods in terms of fidelity and controllability.

可控生成线性注意力边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。