将人脸图像转为可文本控制的矢量素描,提升可编辑性与细节还原度。
PortraVec: Image-Based Portrait Vectorization with Text-Guided Manipulation
- 分两阶段生成:先用注意力偏移采样捕捉人脸结构,再通过区域参数冻结实现局部语义编辑。
- 在多个数据集上对比,结构一致性、视觉保真度和语义可控性均优于现有方法。
- 适合需要精细调整人脸特征的数字艺术创作与交互式设计场景。
尽管人脸素描生成是草图合成中的特殊任务,但现有方法多基于像素,限制了其可解释性和可编辑性。随着矢量生成技术的发展,使用矢量元素表示草图可提供更灵活的操控能力。然而,由于矢量图形的重叠特性及粗粒度细节建模,现有矢量化方法难以保持面部完整性与细粒度细节,且缺乏语义控制。为此,我们提出PortraVec框架,将像素级人脸图像转化为支持文本控制的矢量素描。具体而言,我们设计了一种两阶段图像引导生成模块,采用注意力感知偏移采样以捕捉人脸结构并校正细节偏差;同时提出基于区域参数冻结的文本引导操控模块,实现局部语义编辑的同时保持全局一致性。实验表明,PortraVec在结构一致性、视觉保真度和语义可控性方面均优于当前最优方法。
原文摘要 · Abstract (English)
While portrait sketch generation is a special task in sketch synthesis, most existing methods are pixel-based, limiting their interpretability and editability. With the rise of vector generation techniques, representing sketches using vector elements may provide more flexible manipulation. However, due to the overlapping nature of vector graphics and the coarse detail modeling, existing vectorization methods struggle to capture facial integrity and fine-grained details, and lack semantic control. To address these issues, we propose PortraVec, a framework for converting pixel-based portrait images into vector sketches with text control. Specifically, we propose a two-stage image-guided generation module using Attention-aware Offset Sampling to capture face structure while correcting detail deviations, and a text-guided manipulation module based on Region-based Parameter Freezing to enable local semantic editing while maintaining global consistency. Experiments show that PortraVec achieves superior structural consistency, visual fidelity, and semantic controllability compared to state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。