arXiv:2502.11163cs.CVcs.CL2025-02EMNLP被引 13

AI识图定位有偏,富裕地区准确率高17%以上

AI Sees Your Location, But With A Bias Toward The Wealthy World

  • 构建1200张带地理标签的图像基准测试集
  • 模型在发达地区城市识别准确率达53.8%,欠发达地区低17%
  • 对悉尼等特定城市存在过拟合,隐私风险上升

视觉语言模型(VLMs)在图像地理信息识别任务中表现优异,但存在显著区域偏差。我们构建了一个包含1200张图像及详细地理元数据的基准测试集,评估了四种VLMs。结果显示,尽管模型在城市预测上最高达到53.8%准确率,但在经济欠发达和人口稀疏地区性能分别下降12.5%和17.0%。此外,模型对某些地点如澳大利亚的悉尼存在持续高预测倾向,表现为相关国家熵值偏低。高性能也引发隐私担忧,尤其对无意暴露位置的图像分享者。

原文摘要 · Abstract (English)

Visual-Language Models (VLMs) have shown remarkable performance across various tasks, particularly in recognizing geographic information from images. However, VLMs still show regional biases in this task. To systematically evaluate these issues, we introduce a benchmark consisting of 1,200 images paired with detailed geographic metadata. Evaluating four VLMs, we find that while these models demonstrate the ability to recognize geographic information from images, achieving up to 53.8% accuracy in city prediction, they exhibit significant biases. Specifically, performance is substantially higher for economically developed and densely populated regions compared to less developed (-12.5%) and sparsely populated (-17.0%) areas. Moreover, regional biases of frequently over-predicting certain locations remain. For instance, they consistently predict Sydney for images taken in Australia, shown by the low entropy scores for these countries. The strong performance of VLMs also raises privacy concerns, particularly for users who share images online without the intent of being identified. Our code and dataset are publicly available at https://github.com/uscnlp-lime/FairLocator.

视觉语言模型地理识别偏见检测隐私风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。