用知识图谱嵌入推断用户隐私属性,提出防御方法
Inferring Sensitive Attributes from Knowledge Graph Embeddings: Attack and Defense Strategies
- 从非敏感嵌入输出中逆向推断用户敏感属性
- 随机化后处理可有效降低隐私泄露风险
- 适合关注知识图谱隐私保护的研究者
知识图谱(KGs)是链接数据的强大表示,具有灵活性、语义丰富性,并支持知识增强与推理。尽管如此,现实中的知识图谱常不完整,隐藏真实信息或缺失关键洞察。知识图谱嵌入(KGE)技术常用于推断缺失信息,但基于KGE的推理可能无意间暴露敏感用户信息,即使这些数据未被显式存储。本文研究了基于KGE推理带来的隐私风险,聚焦于属性推断攻击——攻击者试图从看似无敏感性的输出中推断用户的敏感属性。我们提出并评估了一个框架,通过在KGE输出上应用后处理净化技术来缓解此类隐私风险。初步结果表明,该类攻击在多种KGE模型输出上均有效;同时,随机化方法在推荐质量与隐私保护之间存在权衡,提示未来需探索更先进的技术以解决此问题。
原文摘要 · Abstract (English)
Knowledge Graphs (KGs) are a powerful representation of linked data, offering flexibility, semantic richness, and support for knowledge enrichment and reasoning. They help data owners organize and exploit heterogeneous data to provide insightful services (e.g., recommendations), yet real-world KGs are often incomplete, hiding true facts or missing valuable insights. Knowledge graph embedding techniques are commonly used to infer valuable missing information. However, reasoning over KGs can inadvertently expose sensitive user information, even when such data is not explicitly stored. In this work, we investigate the privacy risks associated with KGE-based reasoning, focusing on attribute inference attacks where adversaries attempt to deduce sensitive user attributes from seemingly non-sensitive outputs. We propose and evaluate a framework that mitigates these privacy risks by applying post processing sanitization techniques to KGE outputs. Preliminary results demonstrate the effectiveness of these attacks on the outputs of KGE models, and explore the trade-off between recommendation quality and privacy protection when applying randomization based approaches, highlighting the need to experiment with more advanced techniques in future work to address this issue.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。