用解耦注意力提升文本与面具协同生成人脸的精度与效率
Multivariate Diffusion Transformer with Decoupled Attention for High-Fidelity Mask-Text Collaborative Facial Generation

- 设计统一标记策略与多变量变换器块,同步处理文本与面具信息
- 解耦注意力机制使掩码计算开销降低94%以上,性能不降
- 适合需要高保真人脸生成与高效推理的应用场景
尽管基于语义掩码和文本描述的多模态人脸生成已取得显著进展,但传统特征融合方法难以实现有效的跨模态交互,导致生成效果不佳。为此,我们提出MDiTFace——一种定制化的扩散变换器框架,采用统一标记策略处理语义掩码和文本输入,消除异构模态表示间的差异。通过堆叠新型多变量变换器块,实现所有条件的同步处理与全面多模态特征交互。此外,我们设计了一种解耦注意力机制,将掩码标记与时间嵌入间的隐式依赖关系分离,将内部计算划分为动态与静态路径。静态路径的特征可缓存复用,使掩码条件引入的额外计算开销减少超过94%,同时保持性能。大量实验表明,MDiTFace在人脸保真度与条件一致性方面显著优于现有方法。
原文摘要 · Abstract (English)
While significant progress has been achieved in multimodal facial generation using semantic masks and textual descriptions, conventional feature fusion approaches often fail to enable effective cross-modal interactions, thereby leading to suboptimal generation outcomes. To address this challenge, we introduce MDiTFace--a customized diffusion transformer framework that employs a unified tokenization strategy to process semantic mask and text inputs, eliminating discrepancies between heterogeneous modality representations. The framework facilitates comprehensive multimodal feature interaction through stacked, newly designed multivariate transformer blocks that process all conditions synchronously. Additionally, we design a novel decoupled attention mechanism by dissociating implicit dependencies between mask tokens and temporal embeddings. This mechanism segregates internal computations into dynamic and static pathways, enabling caching and reuse of features computed in static pathways after initial calculation, thereby reducing additional computational overhead introduced by mask condition by over 94% while maintaining performance. Extensive experiments demonstrate that MDiTFace significantly outperforms other competing methods in terms of both facial fidelity and conditional consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。