arXiv:2502.06288cs.CV2025-02中稿 · AI4MFDD 2025 works…被引 3

用语义分割提升地空图像匹配,助力虚假信息检测

Enhancing Ground-to-Aerial Image Matching for Visual Misinformation Detection Using Semantic Segmentation

  • 设计四流孪生网络,融合地表与卫星图像的语义信息
  • 在CVUSA数据集上比现有方法最高提升9.8%匹配准确率
  • 适合媒体可信度审查、数字取证等场景使用

生成式AI技术的快速发展加剧了网络上篡改图像和视频的传播,严重威胁互联网及社交媒体中数字内容的真实性。这一问题在新闻报道、司法鉴定和地球观测等领域尤为突出。为应对挑战,无需外部信息(如GPS坐标)对无地理标签的地表图像进行定位变得愈发关键。本文解决的是在视场角(FoV)变化条件下,将一张地表图像与对应的卫星图像进行匹配的问题。为此,提出一种新型四流孪生结构——四重语义对齐网络(SAN-QUAD),通过在地表与卫星影像上同时应用语义分割,增强特征匹配能力。在CVUSA数据集子集上的实验表明,该方法在多种视场角设置下相比先前最先进方法最高提升9.8%。

原文摘要 · Abstract (English)

The recent advancements in generative AI techniques, which have significantly increased the online dissemination of altered images and videos, have raised serious concerns about the credibility of digital media available on the Internet and distributed through information channels and social networks. This issue particularly affects domains that rely heavily on trustworthy data, such as journalism, forensic analysis, and Earth observation. To address these concerns, the ability to geolocate a non-geo-tagged ground-view image without external information, such as GPS coordinates, has become increasingly critical. This study tackles the challenge of linking a ground-view image, potentially exhibiting varying fields of view (FoV), to its corresponding satellite image without the aid of GPS data. To achieve this, we propose a novel four-stream Siamese-like architecture, the Quadruple Semantic Align Net (SAN-QUAD), which extends previous state-of-the-art (SOTA) approaches by leveraging semantic segmentation applied to both ground and satellite imagery. Experimental results on a subset of the CVUSA dataset demonstrate significant improvements of up to 9.8% over prior methods across various FoV settings.

图像匹配语义分割虚假信息检测地空图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。