用简化矢量图实现图像元素级精确控制,改形状、换颜色一键生成逼真图像。
Controlling Your Image via Simplified Vector Graphics
- 将图像解析为语义对齐的分层矢量图,支持结构化编辑。
- 基于矢量图引导生成,实现几何、色彩、语义的精准调控。
- 适合需要细粒度图像修改的设计师与创作者使用。
图像生成技术虽已达到惊人视觉质量,但核心挑战仍在:能否实现元素级可控?本文提出通过简化矢量图(VGs)实现分层可控生成。方法首先高效地将图像解析为语义对齐、结构连贯的分层矢量表示;在此基础上,设计一种由矢量图引导的新图像合成框架,允许用户自由修改元素,并无缝转化为逼真输出。结合矢量图的结构与语义特征及噪声预测,本方法可精准控制几何形状、颜色和对象语义。大量实验验证了其在图像编辑、对象级操作和细粒度内容创作中的有效性,建立了一种新的可控图像生成范式。
原文摘要 · Abstract (English)
Recent advances in image generation have achieved remarkable visual quality, while a fundamental challenge remains: Can image generation be controlled at the element level, enabling intuitive modifications such as adjusting shapes, altering colors, or adding and removing objects? In this work, we address this challenge by introducing layer-wise controllable generation through simplified vector graphics (VGs). Our approach first efficiently parses images into hierarchical VG representations that are semantic-aligned and structurally coherent. Building on this representation, we design a novel image synthesis framework guided by VGs, allowing users to freely modify elements and seamlessly translate these edits into photorealistic outputs. By leveraging the structural and semantic features of VGs in conjunction with noise prediction, our method provides precise control over geometry, color, and object semantics. Extensive experiments demonstrate the effectiveness of our approach in diverse applications, including image editing, object-level manipulation, and fine-grained content creation, establishing a new paradigm for controllable image generation. Project page: https://guolanqing.github.io/Vec2Pix/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。