用文字干扰保护位置隐私,不破坏图像质量
Beyond Pixels: Semantic-aware Typographic Attack for Geo-Privacy Protection
- 在图像外添加语义误导性文字干扰定位
- 使五款主流大模型定位准确率大幅下降
- 适合关注隐私安全的社交媒体用户
大型视觉语言模型(LVLMs)正构成严重但被忽视的隐私威胁,可直接从分享的图像中推断用户地理位置,导致意外隐私泄露。尽管对抗性图像扰动是潜在防护方向,但需较强畸变才能有效,显著降低图像质量并削弱其传播价值。为此,我们提出一种基于文字扩展的类型学攻击,通过添加语义误导性文本实现隐私保护。进一步研究何种语义内容有效干扰定位,设计出两阶段语义感知型文字攻击,生成欺骗性文字以保护用户隐私。在三个数据集上的大量实验表明,该方法显著降低五款先进商业级LVLM的地理定位准确率,建立了一种实用且保持视觉质量的隐私防护策略。
原文摘要 · Abstract (English)
Large Visual Language Models (LVLMs) now pose a serious yet overlooked privacy threat, as they can infer a social media user's geolocation directly from shared images, leading to unintended privacy leakage. While adversarial image perturbations provide a potential direction for geo-privacy protection, they require relatively strong distortions to be effective against LVLMs, which noticeably degrade visual quality and diminish an image's value for sharing. To overcome this limitation, we identify typographical attacks as a promising direction for protecting geo-privacy by adding text extension outside the visual content. We further investigate which textual semantics are effective in disrupting geolocation inference and design a two-stage, semantics-aware typographical attack that generates deceptive text to protect user privacy. Extensive experiments across three datasets demonstrate that our approach significantly reduces geolocation prediction accuracy of five state-of-the-art commercial LVLMs, establishing a practical and visually-preserving protection strategy against emerging geo-privacy threats.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。