arXiv:2604.24718cs.CV2026-04被引 1

用单目无人机视频生成3D动物检测,无需复杂设备

WildLIFT: Lifting monocular drone video to 3D for species-agnostic wildlife monitoring

论文配图:WildLIFT: Lifting monocular drone video to 3D for species-agnostic wildlife monitoring
图 1 · 摘自论文原文
  • 结合单目视频几何与开放词汇分割,实现无物种依赖的3D检测
  • 在6700+个3D检测上保持高身份一致性,标注效率提升显著
  • 适合生态行为研究与种群监测,输出带视角信息的结构化数据

安装在无人机上的单目RGB相机广泛用于野生动物监测,但多数分析仍局限于二维图像空间,未充分利用视频中的几何信息。我们提出WildLIFT,一种计算框架,将单目无人机视频的三维场景几何与开放词汇2D实例分割相结合,实现无物种依赖的3D检测与跟踪。带有语义面信息的定向3D边界框可量化视角覆盖度与动物间遮挡情况,为下游生态分析提供结构化元数据。我们在2,581帧人工校准数据上验证框架,涵盖四种大型哺乳动物,共超过6,700个3D检测。WildLIFT在多动物场景中保持高身份一致性,并通过关键帧精修大幅减少手动3D标注工作量。通过将标准无人机影像转化为结构化3D与视角感知表示,WildLIFT扩展了航拍野生动物数据集在行为研究与种群监测中的分析潜力。

原文摘要 · Abstract (English)

Monocular RGB cameras mounted on drones are widely used for wildlife monitoring, yet most analytical pipelines remain confined to two-dimensional image space, leaving geometric information in video underexploited. We present WildLIFT, a computational framework that integrates three-dimensional scene geometry from monocular drone video with open-vocabulary 2D instance segmentation to enable species-agnostic 3D detection and tracking. Oriented 3D bounding box labels with semantic face information enable quantitative assessment of viewpoint coverage and inter-animal occlusion, producing structured metadata for downstream ecological analyses. We validate the framework on 2,581 manually curated frames comprising over 6,700 3D detections across four large mammal species. WildLIFT maintains high identity consistency in multi-animal scenes and substantially reduces manual 3D annotation effort through keyframe-based refinement. By transforming standard drone footage into structured 3D and viewpoint-aware representations, WildLIFT extends the analytical utility of aerial wildlife datasets for behavioural research and population monitoring.

3D检测无人机监控生态监测单目视觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。