arXiv:2410.15971cs.CV2024-10NeurIPS被引 40

仅用大模型的先验知识,零样本重建单图场景

Zero-Shot Scene Reconstruction from Single Images with Deep Prior Assembly

  • 融合多个大模型的视觉先验,无需训练即可重建
  • 在多个数据集上优于最新方法,泛化能力强
  • 适合做开放世界场景重建的研究者参考

大规模语言与视觉模型正推动视觉计算的变革。通过大幅扩展数据和模型参数规模,大模型学习到深层先验,显著提升各类任务表现。本文提出深度先验组装(Deep Prior Assembly)框架,通过零样本方式从单张图像中重构场景,仅依赖单一子任务中的深层先验进行泛化。为此,我们设计了关于姿态、尺度和遮挡解析的新方法,使不同先验能稳健协同工作。该框架不需任何3D或2D数据驱动的训练,在开放世界场景中展现出优异的先验泛化能力。我们在多个数据集上进行评估,通过数值、可视化对比及分析,验证了其相对于最新方法的优势。项目主页:https://junshengzhou.github.io/DeepPriorAssembly。

原文摘要 · Abstract (English)

Large language and vision models have been leading a revolution in visual computing. By greatly scaling up sizes of data and model parameters, the large models learn deep priors which lead to remarkable performance in various tasks. In this work, we present deep prior assembly, a novel framework that assembles diverse deep priors from large models for scene reconstruction from single images in a zero-shot manner. We show that this challenging task can be done without extra knowledge but just simply generalizing one deep prior in one sub-task. To this end, we introduce novel methods related to poses, scales, and occlusion parsing which are keys to enable deep priors to work together in a robust way. Deep prior assembly does not require any 3D or 2D data-driven training in the task and demonstrates superior performance in generalizing priors to open-world scenes. We conduct evaluations on various datasets, and report analysis, numerical and visual comparisons with the latest methods to show our superiority. Project page: https://junshengzhou.github.io/DeepPriorAssembly.

场景重建零样本深度先验单图重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。