arXiv:2601.22929cs.CVcs.CL2026-01

压缩图像嵌入中存在语义泄露,可无需重建原图即可还原内容信息。

Semantic Leakage from Image Embeddings

  • 通过保留局部语义邻域结构,从嵌入中恢复语义信息。
  • 在多个模型上验证,均能稳定还原标签、描述等语义内容。
  • 适用于研究隐私安全者,尤其关注嵌入模型风险的开发者。

图像嵌入通常被认为隐私风险有限。本文提出语义泄露概念,即从压缩图像嵌入中恢复语义结构的能力。令人惊讶的是,无需精确重建原始图像,仅需保持嵌入对齐下的局部语义邻域,即可暴露图像嵌入的根本脆弱性。该邻域结构使语义信息可通过一系列有损映射传播。基于此,我们提出SLImE框架,一种轻量级推理方法,利用本地训练的语义检索器与现成模型,无需任务特定解码器,即可从独立压缩嵌入中揭示语义信息。我们在包括GEMINI、COHERE、NOMIC和CLIP在内的多种开放与闭源嵌入模型上全面验证了该框架,涵盖对齐嵌入到标签检索、符号表示及语法连贯描述的全过程。结果表明,语义邻域的保持使语义泄露成为可能,揭示了图像嵌入中的根本隐私风险。

原文摘要 · Abstract (English)

Image embeddings are generally assumed to pose limited privacy risk. We challenge this assumption by formalizing semantic leakage as the ability to recover semantic structures from compressed image embeddings. Surprisingly, we show that semantic leakage does not require exact reconstruction of the original image. Preserving local semantic neighborhoods under embedding alignment is sufficient to expose the intrinsic vulnerability of image embeddings. Crucially, this preserved neighborhood structure allows semantic information to propagate through a sequence of lossy mappings. Based on this conjecture, we propose Semantic Leakage from Image Embeddings (SLImE), a lightweight inference framework that reveals semantic information from standalone compressed image embeddings, incorporating a locally trained semantic retriever with off-the-shelf models, without training task-specific decoders. We thoroughly validate each step of the framework empirically, from aligned embeddings to retrieved tags, symbolic representations, and grammatical and coherent descriptions. We evaluate SLImE across a range of open and closed embedding models, including GEMINI, COHERE, NOMIC, and CLIP, and demonstrate consistent recovery of semantic information across diverse inference tasks. Our results reveal a fundamental vulnerability in image embeddings, whereby the preservation of semantic neighborhoods under alignment enables semantic leakage, highlighting challenges for privacy preservation.1

语义泄露图像嵌入隐私安全模型风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。