arXiv:2503.11633cs.CV2025-03ICCV被引 8

首个支持多层深度估计的图文数据集,助力透明物体三维感知。

Seeing and Seeing Through the Glass: Real and Synthetic Data for Multi-Layer Depth Estimation

  • 构建真实与合成双数据集,支持透明物体多层深度建模。
  • 合成数据训练使模型在真实场景准确率从55.14%提升至75.20%。
  • 适合做透明物体视觉理解、3D重建与机器人感知的研究者。

透明物体在日常生活中普遍存在,理解其多层深度信息——即同时感知透明表面与背后物体——对交互式应用至关重要。本文提出LayeredDepth,首个包含多层深度标注的数据集,包含一个真实世界基准和一个全程序化合成数据生成器,用于支持多层深度估计任务。真实基准由1,500张多样场景图像组成,评估表明现有深度估计方法在透明物体上表现不佳。合成生成器可无限生成不同物体与场景组合,据此构建了含15,300张图像的合成数据集。仅用该合成数据训练的基础模型即可实现跨域多层深度估计;将先进单层深度模型在此数据上微调后,在基准上四倍准确率从55.14%提升至75.20%。所有图像与验证标注均以CC0协议公开,地址为https://layereddepth.cs.princeton.edu。

原文摘要 · Abstract (English)

Transparent objects are common in daily life, and understanding their multi-layer depth information -- perceiving both the transparent surface and the objects behind it -- is crucial for real-world applications that interact with transparent materials. In this paper, we introduce LayeredDepth, the first dataset with multi-layer depth annotations, including a real-world benchmark and a synthetic data generator, to support the task of multi-layer depth estimation. Our real-world benchmark consists of 1,500 images from diverse scenes, and evaluating state-of-the-art depth estimation methods on it reveals that they struggle with transparent objects. The synthetic data generator is fully procedural and capable of providing training data for this task with an unlimited variety of objects and scene compositions. Using this generator, we create a synthetic dataset with 15,300 images. Baseline models training solely on this synthetic dataset produce good cross-domain multi-layer depth estimation. Fine-tuning state-of-the-art single-layer depth models on it substantially improves their performance on transparent objects, with quadruplet accuracy on our benchmark increased from 55.14% to 75.20%. All images and validation annotations are available under CC0 at https://layereddepth.cs.princeton.edu.

深度估计透明物体合成数据多层结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。