让修复人脸既高清又可按指令调整属性,解决模糊与失控难题。
A2BFR: Attribute-Aware Blind Face Restoration
- 用文本提示和退化图像联合控制去噪过程,实现可控修复。
- 在严重退化下仍保持高保真度,属性识别准确率提升52.58%。
- 新数据集支持细粒度属性控制,适合需要精准编辑的场景。
盲人脸修复(BFR)旨在从退化输入中恢复高质量人脸图像,但其固有的病态性导致解不唯一且不可控。现有基于扩散模型的BFR方法虽提升感知质量,却缺乏可控性;而文本引导的人脸编辑虽可操控属性,但修复可靠性不足。为此,本文提出A²BFR框架,统一高保真重建与提示可控生成。基于带有跨模态注意力的扩散变压器骨干网络,A²BFR同时以退化输入和文本提示为条件进行去噪。为引入语义先验,设计属性感知学习,利用属性感知编码器提取的人脸属性嵌入监督去噪潜变量。为进一步增强提示可控性,提出语义双训练策略,借助自建的AttrFace-90K数据集中成对属性变化,强化属性区分性的同时保持修复保真度。大量实验表明,A²BFR在修复保真度与指令遵循性上均达当前最优,相比扩散基线降低0.0467 LPIPS,属性准确率提升52.58%,并在严重退化下仍实现细粒度、提示可控的修复。
原文摘要 · Abstract (English)
Blind face restoration (BFR) aims to recover high-quality facial images from degraded inputs, yet its inherently ill-posed nature leads to ambiguous and uncontrollable solutions. Recent diffusion-based BFR methods improve perceptual quality but remain uncontrollable, whereas text-guided face editing enables attribute manipulation without reliable restoration. To address these issues, we propose A$^2$BFR, an attribute-aware blind face restoration framework that unifies high-fidelity reconstruction with prompt-controllable generation. Built upon a Diffusion Transformer backbone with unified image-text cross-modal attention, A$^2$BFR jointly conditions the denoising trajectory on both degraded inputs and textual prompts. To inject semantic priors, we introduce attribute-aware learning, which supervises denoising latents using facial attribute embeddings extracted by an attribute-aware encoder. To further enhance prompt controllability, we introduce semantic dual-training, which leverages the pairwise attribute variations in our newly curated AttrFace-90K dataset to enforce attribute discrimination while preserving fidelity. Extensive experiments demonstrate that A$^2$BFR achieves state-of-the-art performance in both restoration fidelity and instruction adherence, outperforming diffusion-based BFR baselines by -0.0467 LPIPS and +52.58% attribute accuracy, while enabling fine-grained, prompt-controllable restoration even under severe degradations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。