arXiv:2603.05888cs.CVcs.GR2026-03被引 1

从单张图片生成完整3D室内场景网格,一步完成布局与几何重建。

PixARMesh: Autoregressive Mesh-Native Single-View Scene Reconstruction

  • 基于统一序列的自回归生成,联合预测物体布局与三维几何。
  • 在合成与真实数据集上均达到当前最优重建质量,输出轻量高保真网格。
  • 适合需要快速生成可直接使用的3D场景的应用开发者。

我们提出PixARMesh,一种从单张RGB图像自回归重建完整3D室内场景网格的方法。与依赖隐式符号距离场和后期布局优化的先前方法不同,PixARMesh在统一模型中联合预测物体布局与几何,实现一次前向传播生成连贯且艺术家可用的网格。基于最新网格生成模型进展,我们通过交叉注意力将像素对齐的图像特征与全局场景上下文融入点云编码器,实现从单图准确的空间推理。场景从包含上下文、姿态和网格信息的统一标记序列中自回归生成,产出紧凑且高保真的网格。在合成与真实世界数据集上的实验表明,PixARMesh实现了最先进的重建质量,同时生成适用于下游应用的轻量化高质量网格。

原文摘要 · Abstract (English)

We introduce PixARMesh, a method to autoregressively reconstruct complete 3D indoor scene meshes directly from a single RGB image. Unlike prior methods that rely on implicit signed distance fields and post-hoc layout optimization, PixARMesh jointly predicts object layout and geometry within a unified model, producing coherent and artist-ready meshes in a single forward pass. Building on recent advances in mesh generative models, we augment a point-cloud encoder with pixel-aligned image features and global scene context via cross-attention, enabling accurate spatial reasoning from a single image. Scenes are generated autoregressively from a unified token stream containing context, pose, and mesh, yielding compact meshes with high-fidelity geometry. Experiments on synthetic and real-world datasets show that PixARMesh achieves state-of-the-art reconstruction quality while producing lightweight, high-quality meshes ready for downstream applications.

3D重建自回归生成网格生成单视图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。