构建多模态野生动物三维感知数据集,提升深度估计与重建精度。
WildDepth: A Multimodal Dataset for 3D Wildlife Perception and Depth Estimation
- 同步采集RGB与LiDAR数据,支持多模态深度估计。
- 融合多模态数据使深度误差降低10%,3D重建精度提升12%。
- 适用于野生动物行为分析与跨域泛化感知系统研究。
深度估计与三维重建是计算机视觉的核心课题。研究已从结构简单的刚性物体(如车辆)拓展至复杂可变形物体(如人和动物)。然而,现有动物相关模型大多基于无度量尺度的数据集训练,难以验证仅图像模型的性能。为解决此问题,我们提出WildDepth——一个涵盖家养与野生环境、包含多种动物类别的多模态数据集与基准测试套件,提供同步的RGB与LiDAR数据。实验表明,多模态数据可将深度估计的均方根误差(RMSE)降低10%,而RGB-LiDAR融合使三维重建在Chamfer距离上提升12%。通过发布WildDepth及其基准,我们旨在推动具备跨域泛化能力的鲁棒多模态感知系统发展。
原文摘要 · Abstract (English)
Depth estimation and 3D reconstruction have been extensively studied as core topics in computer vision. Starting from rigid objects with relatively simple geometric shapes, such as vehicles, the research has expanded to address general objects, including challenging deformable objects, such as humans and animals. However, for the animal, in particular, the majority of existing models are trained based on datasets without metric scale, which can help validate image-only models. To address this limitation, we present WildDepth, a multimodal dataset and benchmark suite for depth estimation, behavior detection, and 3D reconstruction from diverse categories of animals ranging from domestic to wild environments with synchronized RGB and LiDAR. Experimental results show that the use of multi-modal data improves depth reliability by up to 10% RMSE, while RGB-LiDAR fusion enhances 3D reconstruction fidelity by 12% in Chamfer distance. By releasing WildDepth and its benchmarks, we aim to foster robust multimodal perception systems that generalize across domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。