arXiv:2506.03371cs.CV2025-06

首个韩语街景细粒度评测基准,揭示视觉语言模型的定位偏见与隐私风险

Toward Reliable VLM: A Fine-Grained Benchmark and Framework for Exposure, Bias, and Inference in Korean Street Views

  • 构建韩国街景细粒度数据集,含多场景标注与真实隐私暴露模拟
  • 10个主流模型在多模态输入下定位精度差异显著,核心城市存在结构性偏差
  • 提出三路径评估框架,支持隐私敏感型视觉语言模型测试

近年来,视觉语言模型(VLMs)在图像地理定位任务中取得显著进展,引发社交媒体日常内容中的位置隐私泄露担忧。然而现有评测基准仍存在粒度粗、语言偏倚、缺乏多模态与隐私感知评估等问题。为此,我们提出 KoreaGEO Bench,首个面向韩国街景的细粒度、多模态地理定位评测基准。数据集包含1,080张高分辨率图像,覆盖四个城市集群和九种场所类型,配备多上下文标注及两种风格的韩语描述,模拟真实世界隐私暴露。我们设计三路径评估协议,测试十款主流VLMs在不同输入模态下的表现,分析其准确性、空间偏差与推理行为。结果表明,输入模态变化导致定位精度显著波动,且模型普遍存在对核心城市的结构性预测偏见。

原文摘要 · Abstract (English)

Recent advances in vision-language models (VLMs) have enabled accurate image-based geolocation, raising serious concerns about location privacy risks in everyday social media posts. However, current benchmarks remain coarse-grained, linguistically biased, and lack multimodal and privacy-aware evaluations. To address these gaps, we present KoreaGEO Bench, the first fine-grained, multimodal geolocation benchmark for Korean street views. Our dataset comprises 1,080 high-resolution images sampled across four urban clusters and nine place types, enriched with multi-contextual annotations and two styles of Korean captions simulating real-world privacy exposure. We introduce a three-path evaluation protocol to assess ten mainstream VLMs under varying input modalities and analyze their accuracy, spatial bias, and reasoning behavior. Results reveal modality-driven shifts in localization precision and highlight structural prediction biases toward core cities.

视觉语言模型地理定位隐私安全多模态评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。