用多视角图像修复低精度体素,提升3D建模质量
VIAFormer: Voxel-Image Alignment Transformer for High-Fidelity Voxel Refinement
- 通过图像索引实现2D图像与3D体素的空间对齐
- 在合成和真实噪声数据上均超越现有方法
- 适合需要高质量3D重建的工业级应用
我们提出VIAFormer,一种用于多视角条件体素精炼的体素-图像对齐变换器,旨在利用校准的多视角图像指导修复不完整、含噪的体素。其有效性源于协同设计:图像索引为2D图像标记提供显式的3D空间定位,修正流目标学习直接的体素修复路径,混合流变换器实现鲁棒的跨模态融合。实验表明,VIAFormer在修复由强大视觉基础模型生成的体素形状上的严重合成噪声和真实伪影方面建立了新基准。除了基准测试,我们还展示了VIAFormer在真实3D创作流程中作为实用可靠桥梁的能力,为基于体素的方法在大模型、大数据浪潮中蓬勃发展铺平道路。
原文摘要 · Abstract (English)
We propose VIAFormer, a Voxel-Image Alignment Transformer model designed for Multi-view Conditioned Voxel Refinement--the task of repairing incomplete noisy voxels using calibrated multi-view images as guidance. Its effectiveness stems from a synergistic design: an Image Index that provides explicit 3D spatial grounding for 2D image tokens, a Correctional Flow objective that learns a direct voxel-refinement trajectory, and a Hybrid Stream Transformer that enables robust cross-modal fusion. Experiments show that VIAFormer establishes a new state of the art in correcting both severe synthetic corruptions and realistic artifacts on the voxel shape obtained from powerful Vision Foundation Models. Beyond benchmarking, we demonstrate VIAFormer as a practical and reliable bridge in real-world 3D creation pipelines, paving the way for voxel-based methods to thrive in large-model, big-data wave.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。