arXiv:2511.08224cs.CVcs.AI2025-11

无需高分辨率图像,用2D表示实现实时单视角3D超分辨率。

2D Representation for Unguided Single-View 3D Super-Resolution in Real-Time

  • 将3D几何信息编码为规则2D图像,直接复用2D超分模型
  • Swin Transformer版达顶尖精度,Vision Mamba版实现实时推理
  • 适合无高分辨率图像可用的真实场景应用

我们提出2Dto3D-SR,一种实时单视角3D超分辨率框架,无需高分辨率RGB引导。该框架将单视角3D数据编码为结构化2D表示,使现有2D图像超分辨率架构可直接应用。通过投影归一化坐标码(PNCC)将可见表面的3D几何信息表示为规则图像,规避了3D点云或RGB引导方法的复杂性。该设计支持轻量高效模型,适用于多种部署环境。我们采用两种实现:基于Swin Transformer的高精度版本和基于Vision Mamba的高效率版本。实验表明,Swin Transformer模型在标准基准上达到当前最优精度,Vision Mamba模型在实时速度下表现良好。这证明了我们的几何引导流程是一种简单、可行且实用的解决方案,尤其适用于无法获取高分辨率RGB数据的实际场景。

原文摘要 · Abstract (English)

We introduce 2Dto3D-SR, a versatile framework for real-time single-view 3D super-resolution that eliminates the need for high-resolution RGB guidance. Our framework encodes 3D data from a single viewpoint into a structured 2D representation, enabling the direct application of existing 2D image super-resolution architectures. We utilize the Projected Normalized Coordinate Code (PNCC) to represent 3D geometry from a visible surface as a regular image, thereby circumventing the complexities of 3D point-based or RGB-guided methods. This design supports lightweight and fast models adaptable to various deployment environments. We evaluate 2Dto3D-SR with two implementations: one using Swin Transformers for high accuracy, and another using Vision Mamba for high efficiency. Experiments show the Swin Transformer model achieves state-of-the-art accuracy on standard benchmarks, while the Vision Mamba model delivers competitive results at real-time speeds. This establishes our geometry-guided pipeline as a surprisingly simple yet viable and practical solution for real-world scenarios, especially where high-resolution RGB data is inaccessible.

3D超分辨率单视角实时处理几何表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。