用注意力机制提升多视角图像3D重建精度,效果领先。
Refine3DNet: Scaling Precision in 3D Object Reconstruction from Multi-View RGB Images using Attention
- 结合CNN与Transformer,用自注意力增强特征表达。
- 单视图重建IOU领先4.2%,多视图表现更优。
- 适合需要高精度3D建模的VR、机器人视觉场景。
从多视角2D RGB图像生成3D模型近年来受到广泛关注,可拓展虚拟现实、机器人视觉及人机交互等技术能力。本文提出一种混合策略,结合卷积神经网络与Transformer,包含带自注意力机制的视觉自编码器和3D精修网络,并采用新型联合训练分离优化(JTSO)算法进行训练。无序输入的编码特征经自注意力层转化为增强特征图,解码为初始3D体素并进一步优化。所提网络可从单张或多张任意视角图像生成3D体素。在ShapeNet数据集上的评估表明,结合JTSO的方法在单视图与多视图3D重建任务中均超越现有最优方法,单视图重建平均交并比(IOU)领先其他模型4.2%。
原文摘要 · Abstract (English)
Generating 3D models from multi-view 2D RGB images has gained significant attention, extending the capabilities of technologies like Virtual Reality, Robotic Vision, and human-machine interaction. In this paper, we introduce a hybrid strategy combining CNNs and transformers, featuring a visual auto-encoder with self-attention mechanisms and a 3D refiner network, trained using a novel Joint Train Separate Optimization (JTSO) algorithm. Encoded features from unordered inputs are transformed into an enhanced feature map by the self-attention layer, decoded into an initial 3D volume, and further refined. Our network generates 3D voxels from single or multiple 2D images from arbitrary viewpoints. Performance evaluations using the ShapeNet datasets show that our approach, combined with JTSO, outperforms state-of-the-art techniques in single and multi-view 3D reconstruction, achieving the highest mean intersection over union (IOU) scores, surpassing other models by 4.2% in single-view reconstruction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。