用照片的地理位置信息控制图像生成,让模型学会不同地点的视觉差异。
GPS as a Control Signal for Image Generation
- 用GPS和文本联合控制扩散模型生成图像
- 生成图像能准确体现城市中不同街区、公园和地标特征
- 利用GPS约束3D重建,提升从2D图像恢复的结构质量
我们证明了照片元数据中的GPS标签可作为图像生成的有效控制信号。通过训练基于GPS的图像生成模型,实现对城市内图像细微变化的精准理解。特别地,我们训练了一个以GPS和文本为条件的扩散模型,生成的图像能够捕捉不同社区、公园和地标独特的外观特征。此外,通过分数蒸馏采样,从二维的GPS到图像模型中提取三维模型,并利用GPS条件约束每个视角下的重建外观。评估表明,我们的GPS条件模型成功学习到基于位置的图像变化规律,且GPS条件有助于提升估计的三维结构质量。
原文摘要 · Abstract (English)
We show that the GPS tags contained in photo metadata provide a useful control signal for image generation. We train GPS-to-image models and use them for tasks that require a fine-grained understanding of how images vary within a city. In particular, we train a diffusion model to generate images conditioned on both GPS and text. The learned model generates images that capture the distinctive appearance of different neighborhoods, parks, and landmarks. We also extract 3D models from 2D GPS-to-image models through score distillation sampling, using GPS conditioning to constrain the appearance of the reconstruction from each viewpoint. Our evaluations suggest that our GPS-conditioned models successfully learn to generate images that vary based on location, and that GPS conditioning improves estimated 3D structure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。