从单目图像中恢复动物3D姿态,为无3D标注数据的检测提供新标签生成方法。
Adding Another Dimension to Image-based Animal Detection

- 用SMAL模型生成动物3D框并投影到2D图像生成标签
- 在Animal3D数据集上实现跨物种和场景的准确3D定位
- 提供可视面度量,适合做单目3D动物检测算法研究
单目图像将动物的3D结构简化为2D投影,现有检测算法仅输出2D边界框,无法反映动物相对于相机的姿态。为构建基于RGB图像的3D动物检测方法,缺乏带标注的数据集;标注过程通常需要3D输入流与RGB数据同步。本文提出一种流程,利用皮肤化多动物线性模型(Skinned Multi Animal Linear, SMAL)估计3D边界框,并通过专用相机位姿优化算法将其投影为2D图像空间中的鲁棒标签。计算立方体面可见性指标以判断动物哪些侧面被捕捉。这些3D边界框和可见性指标为未来单目3D动物检测算法的研发与评估提供了关键基础。我们在Animal3D数据集上评估了该方法,在不同物种和场景下均展现出良好性能。
原文摘要 · Abstract (English)
Monocular imaging of animals inherently reduces 3D structures to 2D projections. Detection algorithms lead to 2D bounding boxes that lack information about animal's orientation relative to the camera. To build 3D detection methods for RGB animal images, there is a lack of labeled datasets; such labeling processes require 3D input streams along with RGB data. We present a pipeline that utilises Skinned Multi Animal Linear models to estimate 3D bounding boxes and to project them as robust labels into 2D image space using a dedicated camera pose refinement algorithm. To assess which sides of the animal are captured, cuboid face visibility metrics are computed. These 3D bounding boxes and metrics form a crucial step toward developing and benchmarking future monocular 3D animal detection algorithms. We evaluate our method on the Animal3D dataset, demonstrating accurate performance across species and settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。