统一视角下单目图像生成,适配广角到全景多种镜头
UniSHARP: Universal Sharp Monocular View Synthesis

- 用统一全景隐空间对齐不同镜头图像,实现跨镜头渲染
- 在多类镜头场景下超越现有方法,尤其在大视场表现更优
- 适合需要多镜头兼容的3D内容生成研究与应用
本文致力于将流行的逼真单目视角生成方法SHARP拓展至从标准透视相机到广角、鱼眼及全向全景等多种镜头系统的通用单目渲染。为克服SHARP依赖针孔相机假设的局限,核心思想是将各类图像对齐至统一的全景隐空间。为此提出UniSHARP,通过特征与高斯空间的联合隐式对齐实现。具体而言,高斯原型沿射线和径向距离构建射线基础的通用表示,同时结合源自UniK3D启发的编码器提取的2D语义与3D空间特征,联合解码生成完整高斯云。为全面评估方法,构建覆盖多类成像系统与场景的基准数据集,并按视场角(FoV)分层以实现细粒度评估。在该基准上的大量实验表明,UniSHARP显著优于现有方法。
原文摘要 · Abstract (English)
In this work, we focus on extending SHARP, the popular photorealistic view synthesis method, for universal monocular rendering across a continuum of camera systems, from conventional perspective cameras to wide-field-of-view, fisheye and omnidirectional panoramic settings. To overcome the pinhole-specific assumptions of SHARP, our key idea is to align various images in a unified omnidirectional latent space. Thus, we propose UniSHARP, which performs implicit alignment in both feature and Gaussian spaces. Specifically, Gaussian primitives are arranged along rays and radial distances in a ray-based universal representation, while 2D semantic and 3D spatial features extracted from UniK3D-inspired encoders are jointly decoded to generate the complete Gaussian cloud. To comprehensively evaluate our method, we construct a benchmark covering diverse imaging systems across various scenes. The benchmark is further stratified by field of view (FoV) to enable fine-grained assessment of the universal monocular rendering task. Extensive experiments on the proposed benchmark demonstrate the effectiveness of UniSHARP, outperforming alternative methods by a large margin. The project page can be found at: https://insta360-research-team.github.io/Unisharp-website/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。