arXiv:2512.07834cs.CV2025-12被引 3

让3D模型自动生成像素风立体艺术,精准对齐且色彩可控。

Voxify3D: Pixel Art Meets Volumetric Rendering

  • 分两阶段优化:先调整网格,再用2D像素监督对齐
  • 生成效果得分37.12(CLIP-IQA),用户偏好达77.90%
  • 支持2-8种颜色、20x-50x分辨率的可控抽象

体素艺术是一种广泛应用于游戏与数字媒体的独特风格,但其从3D网格自动生成仍面临几何抽象、语义保留与离散色彩一致性之间的冲突。现有方法或过度简化几何结构,或无法实现像素级、调色板受限的视觉效果。我们提出Voxify3D,一种可微分的两阶段框架,将3D网格优化与2D像素艺术监督相结合。核心创新包括:(1) 正交投影下的像素艺术监督,消除透视失真以实现精确的体素-像素对齐;(2) 基于图像块的CLIP对齐,保持离散化过程中的语义一致性;(3) 调色板约束的Gumbel-Softmax量化,实现离散色彩空间的可微优化,支持可控调色板策略。该框架有效解决极端离散化下的语义保留、像素艺术美学与端到端离散优化等根本挑战。实验表明,在多种角色上表现优异(CLIP-IQA 37.12,用户偏好77.90%),支持2-8种颜色、20x-50x分辨率的可控抽象。

原文摘要 · Abstract (English)

Voxel art is a distinctive stylization widely used in games and digital media, yet automated generation from 3D meshes remains challenging due to conflicting requirements of geometric abstraction, semantic preservation, and discrete color coherence. Existing methods either over-simplify geometry or fail to achieve the pixel-precise, palette-constrained aesthetics of voxel art. We introduce Voxify3D, a differentiable two-stage framework bridging 3D mesh optimization with 2D pixel art supervision. Our core innovation lies in the synergistic integration of three components: (1) orthographic pixel art supervision that eliminates perspective distortion for precise voxel-pixel alignment; (2) patch-based CLIP alignment that preserves semantics across discretization levels; (3) palette-constrained Gumbel-Softmax quantization enabling differentiable optimization over discrete color spaces with controllable palette strategies. This integration addresses fundamental challenges: semantic preservation under extreme discretization, pixel-art aesthetics through volumetric rendering, and end-to-end discrete optimization. Experiments show superior performance (37.12 CLIP-IQA, 77.90% user preference) across diverse characters and controllable abstraction (2-8 colors, 20x-50x resolutions). Project page: https://yichuanh.github.io/Voxify-3D/

体素艺术风格生成可微优化3D生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。