通过空间流分解实现精准文本驱动人脸编辑
Flux-Sculptor: Text-Driven Rich-Attribute Portrait Editing through Decomposed Spatial Flow Control
- 用提示对齐定位器精确定位需编辑区域
- 序列掩码融合潜空间与注意力值,提升编辑精度
- 适合需要高保真人脸保留的创意设计场景
文本驱动的人脸编辑在多种应用中潜力巨大,但面临定位精度与内容修改灵活性难以兼顾的挑战。现有方法往往在重建保真度与编辑灵活性之间权衡困难。为此,我们提出Flux-Sculptor,一种基于通量的框架,实现精确的文本驱动人脸编辑。该框架引入提示对齐空间定位器(PASL),准确识别需编辑区域;并采用结构到细节编辑控制(S2D-EC)策略,通过潜空间表示与注意力值的序列掩码引导融合,空间化地指导去噪过程。大量实验表明,Flux-Sculptor在丰富属性编辑与面部信息保留方面均优于现有方法,是实际人脸编辑应用的有力候选方案。项目页面见https://flux-sculptor.github.io/。
原文摘要 · Abstract (English)
Text-driven portrait editing holds significant potential for various applications but also presents considerable challenges. An ideal text-driven portrait editing approach should achieve precise localization and appropriate content modification, yet existing methods struggle to balance reconstruction fidelity and editing flexibility. To address this issue, we propose Flux-Sculptor, a flux-based framework designed for precise text-driven portrait editing. Our framework introduces a Prompt-Aligned Spatial Locator (PASL) to accurately identify relevant editing regions and a Structure-to-Detail Edit Control (S2D-EC) strategy to spatially guide the denoising process through sequential mask-guided fusion of latent representations and attention values. Extensive experiments demonstrate that Flux-Sculptor surpasses existing methods in rich-attribute editing and facial information preservation, making it a strong candidate for practical portrait editing applications. Project page is available at https://flux-sculptor.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。