arXiv:2606.23669cs.CV2026-06

测试文字生成街景图对具体路段的准确还原能力

GeoFidelity-Bench: Evaluating Segment-Level Geographic Fidelity in Text-to-Image Street-View Generation

论文配图:GeoFidelity-Bench: Evaluating Segment-Level Geographic Fidelity in Text-to-Image Street-View Generation
图 1 · 摘自论文原文
  • 构建基准数据集,对比生成图像与真实路段匹配度
  • 加街道名可提升5.5%识别准确率,但难精准定位路段
  • 仅用城市名生成更泛化,地理坐标文本无额外帮助

文本到图像模型能生成视觉上可信的城市街景,但其输出是否对应指定道路段而非通用城市特征尚不明确。本文提出GeoFidelity-Bench,一个用于评估街景生成中段级地理保真度的参考面板基准。该数据集包含7,117张精心筛选的Mapillary图像,覆盖六大洲25个城市的109条知名OpenStreetMap道路段。针对每张生成图像,基准将其目标参考图像与同城市最近路段、其他路段及异城市路段进行排名,以局部区分性而非绝对相似性为主要测试标准。评估六款开源文本到图像生成器在仅城市名、街道与街区名、以及带GPS提示下的表现。结果显示,加入街道和街区名称使顶1检索准确率比仅城市名提升5.5个百分点(95%置信区间3.4–7.7)。然而,目标图像与同城市最近路段间的相似度差距接近零,表明局部名称更多提升整体本地合理性而非精确段落身份。使用错误街道或街区名的提示显示,部分性能提升并非依赖正确名称;将原始GPS坐标作为普通文本添加亦未带来统计显著优势。保留的真实图像查询可成功恢复路段身份,证明参考数据集包含可用的段级信号。因此,GeoFidelity-Bench揭示了当前城市或街区合理街景生成与特定路段忠实生成之间存在持续差距。

原文摘要 · Abstract (English)

Text-to-image models can generate visually plausible city streets, but whether their outputs correspond to a requested road segment rather than a generic city prior remains unclear. We introduce GeoFidelity-Bench, a reference-panel benchmark for segment-conditioned geographic fidelity in street-view generation. It contains 7,117 curated Mapillary images covering 109 named OpenStreetMap road segments in 25 cities across six continents. For each generated panel, the benchmark ranks the target reference panel against panels from the nearest segment in the same city, other segments in the same city, and segments from other cities, making local discrimination rather than absolute target similarity the primary test. We evaluate six open-weight text-to-image generators under city-only, street-and-neighborhood, and GPS-augmented prompts. Adding street and neighborhood names is associated with an increase of 5.5 percentage points in top-1 retrieval accuracy over city-only prompts, with a 95% confidence interval from 3.4 to 7.7 percentage points. However, the similarity margin between the target and the nearest segment in the same city remains near zero, indicating that local names improve broad local plausibility more than exact segment identity. Prompts that keep the city fixed but use incorrect street or neighborhood names further show that only part of the gain depends on the correct local names, while appending raw GPS coordinates as ordinary text yields no statistically clear additional benefit. Held-out real-image queries successfully recover segment identity, showing that the curated references contain usable segment-level signal. GeoFidelity-Bench thus reveals a persistent gap between city- or neighborhood-plausible street-view generation and faithful generation for a specific road segment.

图像生成地理保真度基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。