arXiv:2503.12464cs.CVcs.LG2025-03中稿 · and to appear in t…被引 2

用简单CNN+迁移学习预测图像隐私,参数少但效果不输复杂图模型。

Learning Privacy from Visual Entities

  • 用迁移学习+轻量CNN关联场景类型与隐私,仅优化732个参数。
  • 相比图神经网络方法,性能相当但模型更简洁。
  • 发现高维特征和图结构对隐私判断贡献不大,适合快速部署。

主观理解与内容多样性使得判断图像是否私密极具挑战。本文采用结合图神经网络与卷积神经网络(CNN)的方法,其参数量在1400万至5000万之间,用于提取视觉实体(如场景、物体类型)的特征,并识别影响判断的关键实体。然而,我们发现仅使用迁移学习与轻量级CNN,通过关联场景类型与隐私,仅需优化732个参数即可达到与图模型相当的性能。相比之下,端到端训练的图模型会掩盖各组件对分类结果的实际贡献。此外,为每个视觉实体提取高维特征既非必要,又增加模型复杂度;图结构本身对性能影响微乎其微,真正关键在于微调CNN以优化隐私相关特征。

原文摘要 · Abstract (English)

Subjective interpretation and content diversity make predicting whether an image is private or public a challenging task. Graph neural networks combined with convolutional neural networks (CNNs), which consist of 14,000 to 500 millions parameters, generate features for visual entities (e.g., scene and object types) and identify the entities that contribute to the decision. In this paper, we show that using a simpler combination of transfer learning and a CNN to relate privacy with scene types optimises only 732 parameters while achieving comparable performance to that of graph-based methods. On the contrary, end-to-end training of graph-based methods can mask the contribution of individual components to the classification performance. Furthermore, we show that a high-dimensional feature vector, extracted with CNNs for each visual entity, is unnecessary and complexifies the model. The graph component has also negligible impact on performance, which is driven by fine-tuning the CNN to optimise image features for privacy nodes.

图像隐私轻量模型迁移学习CNN

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。