用GAN自动生成可触达地图,支持跨尺度跨区域泛化。
A Step towards Automated and Generalizable Tactile Map Generation using Generative Adversarial Networks
- 基于生成对抗网络,从谷歌街景图像中识别关键地图元素。
- 单视图训练模型在所有特征上F1和交并比均超0.97。
- 模型能泛化到未见地图尺度和地理区域,适合无障碍导航应用。
全球有大量盲人和视觉障碍者。为辅助导航,他们常依赖通过触摸感知的触觉地图,利用凸起表面与边缘传递信息。然而这些地图难以普及,现有自动化工具也存在仅适用于特定比例尺、地理区域或触觉制图标准的问题。为此,我们提出首个概念验证模型,首次将计算机视觉技术用于触觉地图自动生成。构建了首个涵盖6500个地点的谷歌街景触觉地图数据集,包含多种线状与面状触觉特征。在单一缩放级别上训练的生成对抗网络(GAN)模型能有效识别关键地图要素、剔除冗余信息,并完成补全,各类特征的平均F1与交并比(IoU)均超过0.97。在双缩放级别训练的模型性能略有下降,但仍表现出良好的跨尺度与跨区域泛化能力。最后讨论了未来实现完整触觉地图解决方案的方向。
原文摘要 · Abstract (English)
Blindness and visual impairments affect many people worldwide. For help with navigation, people with visual impairments often rely on tactile maps that utilize raised surfaces and edges to convey information through touch. Although these maps are helpful, they are often not widely available and current tools to automate their production have similar limitations including only working at certain scales, for particular world regions, or adhering to specific tactile map standards. To address these shortcomings, we train a proof-of-concept model as a first step towards applying computer vision techniques to help automate the generation of tactile maps. We create a first-of-its-kind tactile maps dataset of street-views from Google Maps spanning 6500 locations and including different tactile line- and area-like features. Generative adversarial network (GAN) models trained on a single zoom successfully identify key map elements, remove extraneous ones, and perform inpainting with median F1 and intersection-over-union (IoU) scores of better than 0.97 across all features. Models trained on two zooms experience only minor drops in performance, and generalize well both to unseen map scales and world regions. Finally, we discuss future directions towards a full implementation of a tactile map solution that builds on our results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。