用多视角图像快速精准估算三维体积与表面积,适合数据少或设备弱的场景。
Lightweight Neural Framework for Robust 3D Volume and Surface Estimation from Multi-View Images

- 直接从多视角图像回归归一化体积和表面积及其不确定性。
- 在仅用少量图像时性能优于现有方法,推理速度快且可扩展。
- 适用于珊瑚监测、饮食分析等实际场景,对噪声和稀疏数据鲁棒。
精确估计体积和表面积在海洋生态、医疗诊断等多个领域至关重要。然而,现有方法常面临高计算成本及在稀疏、噪声数据下表现不佳的问题。本文提出一种全前馈框架,直接从多视角图像回归尺度归一化的体积、表面积及其不确定性。通过图结构解码器融合3D点云重建与视图对齐的2D特征,模型避免了迭代优化,实现极佳的可扩展性与快速推理。实验表明,该方法在输入图像数量较少时显著优于当前最优方法。在珊瑚监测、饮食分析和人体测量任务中验证了其鲁棒性与适应性。该架构为视觉数据下的几何估计提供了高速、可扩展的解决方案,在资源受限或视角稀疏场景下仍保持高性能。
原文摘要 · Abstract (English)
Accurate volume and surface area estimation is critical for diverse applications, from marine ecology to medical diagnostics. However, existing methods often suffer from high computational costs and poor performance with sparse and noisy data. We propose a fully feed-forward framework that regresses scale-normalized volume and surface area and their associated uncertainties directly from multi-view images. By fusing 3D point cloud reconstructions with view-aligned 2D features through a graph-based decoder, our model bypasses iterative optimization, ensuring exceptional scalability and rapid inference. Experimental results demonstrate that our approach outperforms state-of-the-art methods, particularly when operating with a low number of input images. Validated across coral monitoring, dietary analysis, and anthropometry, our proposed framework provides a robust, adaptable solution for quantitative shape analysis. This architecture provides a high-speed, scalable alternative for precise geometric estimation from visual data, maintaining high performance even in resource-constrained or sparse-view scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。