arXiv:2409.13535cs.CV2024-09ECCV被引 4

用数学公式自动生成图像与点云,统一训练视觉几何模型。

Formula-Supervised Visual-Geometric Pre-training

论文配图:Formula-Supervised Visual-Geometric Pre-training
图 1 · 摘自论文原文
  • 从数学公式生成配对的图像和点云,实现跨模态监督。
  • 在6项任务中超越现有方法,提升3D识别性能。
  • 减少真实数据依赖,适合多模态感知研究者。

计算机视觉中,图像(视觉)与点云(几何)常被分别处理。本文提出公式监督的视觉-几何预训练(FSVGP),通过数学公式自动生成对齐的合成图像与点云,实现视觉与几何模态间的监督预训练。该方法降低对真实数据、跨模态对齐和人工标注的依赖。实验表明,FSVGP在图像与3D物体分类、检测、分割共六项任务中优于VisualAtom和PC-FractalDB,展现更强的泛化能力,验证了合成预训练在视觉-几何表示学习中的潜力。

原文摘要 · Abstract (English)

Throughout the history of computer vision, while research has explored the integration of images (visual) and point clouds (geometric), many advancements in image and 3D object recognition have tended to process these modalities separately. We aim to bridge this divide by integrating images and point clouds on a unified transformer model. This approach integrates the modality-specific properties of images and point clouds and achieves fundamental downstream tasks in image and 3D object recognition on a unified transformer model by learning visual-geometric representations. In this work, we introduce Formula-Supervised Visual-Geometric Pre-training (FSVGP), a novel synthetic pre-training method that automatically generates aligned synthetic images and point clouds from mathematical formulas. Through cross-modality supervision, we enable supervised pre-training between visual and geometric modalities. FSVGP also reduces reliance on real data collection, cross-modality alignment, and human annotation. Our experimental results show that FSVGP pre-trains more effectively than VisualAtom and PC-FractalDB across six tasks: image and 3D object classification, detection, and segmentation. These achievements demonstrate FSVGP's superior generalization in image and 3D object recognition and underscore the potential of synthetic pre-training in visual-geometric representation learning. Our project website is available at https://ryosuke-yamada.github.io/fdsl-fsvgp/.

视觉几何合成数据多模态预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。