用手机拍的普通照片,也能生成超高清全景图像。
UltraZoom: Generating Gigapixel Images from Regular Photos
- 用局部特写图对齐全局照片,学习物体专属的超分映射。
- 可生成分辨率高达十亿像素的连续高清图像。
- 适合需要高精度图像重建的摄影与考古领域。
我们提出UltraZoom,一个从随意拍摄的输入图像(如手持手机照片)生成物体级十亿像素分辨率图像的系统。给定一张全幅图(全局、低细节)和一张或多张局部特写图(局部、高细节),UltraZoom将全幅图放大至与特写图一致的精细程度和尺度。为此,我们从特写图构建针对特定实例的成对数据集,并微调预训练生成模型以学习对象特有的低-高分辨率映射。推理时,模型以滑动窗口方式作用于全幅图。配对构建非同小可:需在全图中对齐特写图以进行尺度估计和退化对齐。我们提出一种简单而鲁棒的方法,适用于任意材质的随意野外拍摄。这些组件共同构成一个系统,可在整个物体上实现无缝缩放与平移,仅凭最少输入即生成一致且逼真的十亿像素图像。
原文摘要 · Abstract (English)
We present UltraZoom, a system for generating gigapixel-resolution images of objects from casually captured inputs, such as handheld phone photos. Given a full-shot image (global, low-detail) and one or more close-ups (local, high-detail), UltraZoom upscales the full image to match the fine detail and scale of the close-up examples. To achieve this, we construct a per-instance paired dataset from the close-ups and adapt a pretrained generative model to learn object-specific low-to-high resolution mappings. At inference, we apply the model in a sliding window fashion over the full image. Constructing these pairs is non-trivial: it requires registering the close-ups within the full image for scale estimation and degradation alignment. We introduce a simple, robust method for getting registration on arbitrary materials in casual, in-the-wild captures. Together, these components form a system that enables seamless pan and zoom across the entire object, producing consistent, photorealistic gigapixel imagery from minimal input.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。