用高程图生成地形纹理,保持地形与视觉一致
Geodiffussr: Generative Terrain Texturing with Elevation Fidelity
- 通过多尺度特征注入,让纹理贴图严格匹配给定高程图
- 相比基线,图像质量提升49%(FID),高度与外观关联度接近0
- 适合快速构思大规模景观,补充物理模拟工具
大规模地形生成在计算机图形学中仍属劳动密集型任务。我们提出Geodiffussr,一种基于流匹配的生成管道,可在严格遵循输入数字高程图(DEM)的前提下,合成文本引导的纹理贴图。核心机制为多尺度内容聚合(MCA):将预训练编码器提取的DEM特征注入UNet多个分辨率层级,以确保全局到局部的高程一致性。相较于无MCA的基线模型,MCA显著提升视觉保真度,并强化高度与外观的耦合关系(FID下降49.16%,LPIPS下降32.33%,ΔdCor降至0.0016)。为训练与评估,我们构建了一个全球分布、按生物群落和气候分层的三元组数据集,包含由SRTM生成的DEM、Sentinel-2影像及描述可见地表覆盖的视觉语言标注。我们将Geodiffussr定位为粗粒度构想与前期预览的强基准,推动可控2.5D景观生成,与物理基础的地形与生态系统模拟器形成互补。
原文摘要 · Abstract (English)
Large-scale terrain generation remains a labor-intensive task in computer graphics. We introduce Geodiffussr, a flow-matching pipeline that synthesizes text-guided texture maps while strictly adhering to a supplied Digital Elevation Map (DEM). The core mechanism is multi-scale content aggregation (MCA): DEM features from a pretrained encoder are injected into UNet blocks at multiple resolutions to enforce global-to-local elevation consistency. Compared with a non-MCA baseline, MCA markedly improves visual fidelity and strengthens height-appearance coupling (FID $\downarrow$ 49.16%, LPIPS $\downarrow$ 32.33%, $Δ$dCor $\downarrow$ to 0.0016). To train and evaluate Geodiffussr, we assemble a globally distributed, biome- and climate-stratified corpus of triplets pairing SRTM-derived DEMs with Sentinel-2 imagery and vision-grounded natural-language captions that describe visible land cover. We position Geodiffussr as a strong baseline and step toward controllable 2.5D landscape generation for coarse-scale ideation and previz, complementary to physically based terrain and ecosystem simulators.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。