提升物体边界与细节区域的深度估计精度。
Detail-aware multi-view stereo network for depth estimation
- 利用粗到细框架捕捉几何线索,增强特征表达能力。
- 图像合成损失约束梯度流,强化边缘与纹理区监督。
- 自适应调整深度间隔,改善物体重建准确率。
多视图立体方法在基于粗到细深度学习框架下已取得显著进展,但现有方法在恢复物体边界和细节区域深度时表现不佳。为此,我们提出一种细节感知的多视图立体网络(DA-MVSNet),采用粗到细框架。通过利用粗阶段隐藏的几何深度线索,保持物体表面间的几何结构关系,并增强图像特征的表达能力。此外,引入图像合成损失以约束细节区域的梯度流动,进一步加强物体边界和纹理丰富区域的监督。最后,提出自适应深度间隔调整策略,提升物体重建精度。在DTU和Tanks & Temples数据集上的大量实验表明,该方法达到具有竞争力的结果。代码已公开于https://github.com/wsmtht520-/DAMVSNet。
原文摘要 · Abstract (English)
Multi-view stereo methods have achieved great success for depth estimation based on the coarse-to-fine depth learning frameworks, however, the existing methods perform poorly in recovering the depth of object boundaries and detail regions. To address these issues, we propose a detail-aware multi-view stereo network (DA-MVSNet) with a coarse-to-fine framework. The geometric depth clues hidden in the coarse stage are utilized to maintain the geometric structural relationships between object surfaces and enhance the expressive capability of image features. In addition, an image synthesis loss is employed to constrain the gradient flow for detailed regions and further strengthen the supervision of object boundaries and texture-rich areas. Finally, we propose an adaptive depth interval adjustment strategy to improve the accuracy of object reconstruction. Extensive experiments on the DTU and Tanks & Temples datasets demonstrate that our method achieves competitive results. The code is available at https://github.com/wsmtht520-/DAMVSNet.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。