28个全球城市街景数据存在布局偏差,影响视觉模型训练公平性。
Artifacts of Idiosyncracy in Global Street View Data
- 分析28城街景数据覆盖分布,发现城市布局导致采样偏差。
- 提出评估方法,量化不同城市间数据覆盖差异。
- 通过阿姆斯特丹案例研究,揭示采集过程如何引入偏见。
近年来,街景数据在计算机视觉应用中日益普及。机器学习数据集通常采用简单采样方式构建,假设其能系统性代表城市,尤其在密集采样下。然而,已有研究指出某些城市或区域覆盖率显著不足。本文揭示,即使在高密度采样下,28个全球城市的独特布局等特征仍会导致街景数据覆盖偏差。我们定量分析了数据覆盖分布的不均衡性,并提出一种评估方法以深入理解城市层面的覆盖异质性。此外,针对阿姆斯特丹开展半结构化访谈,探讨了采集过程中的特异性如何影响城市与区域的表征,从而为从源头缓解偏见提供依据。
原文摘要 · Abstract (English)
Street view data is increasingly being used in computer vision applications in recent years. Machine learning datasets are collected for these applications using simple sampling techniques. These datasets are assumed to be a systematic representation of cities, especially when densely sampled. Prior works however, show that there are clear gaps in coverage, with certain cities or regions being covered poorly or not at all. Here we demonstrate that a cities' idiosyncracies, such as city layout, may lead to biases in street view data for 28 cities across the globe, even when they are densely covered. We quantitatively uncover biases in the distribution of coverage of street view data and propose a method for evaluation of such distributions to get better insight in idiosyncracies in a cities' coverage. In addition, we perform a case study of Amsterdam with semi-structured interviews, showing how idiosyncracies of the collection process impact representation of cities and regions and allowing us to address biases at their source.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。