arXiv:2411.18807cs.CVcs.CL2024-11CVPR被引 8

从单张图像重建动物与自然环境的三维场景,兼顾生态上下文。

Reconstructing Animals and the Wild

  • 基于大语言模型先验,用自回归模型从CLIP嵌入生成结构化场景表示
  • 在百万级合成数据上训练,可直接泛化到真实图像的动物-环境联合重建
  • 首次实现动物与野生环境一体化重建,适合场景理解与生物行为研究

三维重建作为场景理解的基础,在计算机视觉中至关重要。从二维视觉观测重建三维场景需依赖强先验以消除歧义。现有工作多聚焦于人类中心场景,其特征为平滑表面、一致法向与规则边缘,便于融入强几何归纳偏置。本文关注更复杂挑战:重建包含树木、灌木、岩石和动物的自然场景。尽管已有诸多研究尝试重建野外动物,但均仅关注动物本身,忽视环境上下文,导致分析任务中信息丢失。为此,我们提出一种从单图重建自然场景的方法,基于大语言模型蕴含的世界先验,训练自回归模型将CLIP嵌入解码为包含动物与自然环境(RAW)的结构化组合场景表示。为支持此方法,我们构建了一个包含一百万张图像和数千个资产的合成数据集。所提方法仅在合成数据上训练,却能有效泛化至真实图像中的动物及其环境重建。相关数据集与代码已公开,网址为 https://raw.is.tue.mpg.de/

原文摘要 · Abstract (English)

The idea of 3D reconstruction as scene understanding is foundational in computer vision. Reconstructing 3D scenes from 2D visual observations requires strong priors to disambiguate structure. Much work has been focused on the anthropocentric, which, characterized by smooth surfaces, coherent normals, and regular edges, allows for the integration of strong geometric inductive biases. Here, we consider a more challenging problem where such assumptions do not hold: the reconstruction of natural scenes containing trees, bushes, boulders, and animals. While numerous works have attempted to tackle the problem of reconstructing animals in the wild, they have focused solely on the animal, neglecting environmental context. This limits their usefulness for analysis tasks, as animals exist inherently within the 3D world, and information is lost when environmental factors are disregarded. We propose a method to reconstruct natural scenes from single images. We base our approach on recent advances leveraging the strong world priors ingrained in Large Language Models and train an autoregressive model to decode a CLIP embedding into a structured compositional scene representation, encompassing both animals and the wild (RAW). To enable this, we propose a synthetic dataset comprising one million images and thousands of assets. Our approach, having been trained solely on synthetic data, generalizes to the task of reconstructing animals and their environments in real-world images. We will release our dataset and code to encourage future research at https://raw.is.tue.mpg.de/

三维重建自然场景动物重建合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。