arXiv:2605.10922cs.CV2026-05International Conf…被引 4

让3D生成与输入图像像素对齐,提升真实感。

Pixal3D: Pixel-Aligned 3D Generation from Images

论文配图:Pixal3D: Pixel-Aligned 3D Generation from Images
图 1 · 摘自论文原文
  • 直接在输入视角生成3D,通过像素反投影建立精准对应
  • 在单图和多图场景下显著提升3D生成的像素级保真度
  • 适合需要高保真3D物体或场景生成的研究与应用

近期3D生成模型虽大幅提升了图像到3D合成的质量,实现了更高分辨率的几何结构和更真实的外观,但像素级保真度仍是核心瓶颈。我们指出问题源于隐式的2D-3D对应模糊:多数3D原生生成器在规范空间中合成形状,并通过注意力机制注入图像信息,导致像素与3D点的对应关系不明确。为此,我们借鉴3D重建思想,提出Pixal3D——一种像素对齐的3D生成范式,实现从图像生成高保真3D资产。不同于在规范姿态下生成,Pixal3D直接在输入视角下生成3D,引入像素反投影条件机制,将多尺度图像特征显式映射至3D特征体,建立无歧义的像素-3D对应。实验表明,Pixal3D不仅可扩展、生成高质量3D资产,且显著提升保真度,接近重建水平。此外,通过跨视角聚合反投影特征体,自然拓展至多视角生成;并展示像素对齐生成对场景合成的益处,提出模块化流水线,从图像生成高保真、物体分离的3D场景。Pixal3D首次在大规模上实现3D原生像素对齐生成,为从单图或多图生成高保真3D物体或场景提供了新思路。

原文摘要 · Abstract (English)

Recent advances in 3D generative models have rapidly improved image-to-3D synthesis quality, enabling higher-resolution geometry and more realistic appearance. Yet fidelity, which measures pixel-level faithfulness of the generated 3D asset to the input image, still remains a central bottleneck. We argue this stems from an implicit 2D-3D correspondence issue: most 3D-native generators synthesize shape in canonical space and inject image cues via attention, leaving pixel-to-3D associations ambiguous. To tackle this issue, we draw inspiration from 3D reconstruction and propose Pixal3D, a pixel-aligned 3D generation paradigm for high-fidelity 3D asset creation from images. Instead of generating in a canonical pose, Pixal3D directly generates 3D in a pixel-aligned way, consistent with the input view. To enable this, we introduce a pixel back-projection conditioning scheme that explicitly lifts multi-scale image features into a 3D feature volume, establishing direct pixel-to-3D correspondence without ambiguity. We show that Pixal3D is not only scalable and capable of producing high-quality 3D assets, but also substantially improves fidelity, approaching the fidelity level of reconstruction. Furthermore, Pixal3D naturally extends to multi-view generation by aggregating back-projected feature volumes across views. Finally, we show pixel-aligned generation benefits scene synthesis, and present a modular pipeline that produces high-fidelity, object-separated 3D scenes from images. Pixal3D for the first time demonstrates 3D-native pixel-aligned generation at scale, and provides a new inspiring way towards high-fidelity 3D generation of object or scene from single or multi-view images. Project page: https://ldyang694.github.io/projects/pixal3d/

3D生成像素对齐图像生成多视图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。