arXiv:2607.16280cs.CV2026-07

用微小扰动保护3D人脸隐私,让AI误判面部属性。

3D FaceShell: Attribute Transfer in 3D Face Avatars as a VLM Defense Mechanism

论文配图:3D FaceShell: Attribute Transfer in 3D Face Avatars as a VLM Defense Mechanism
图 1 · 摘自论文原文
  • 在3D人脸模型外加可学习高斯壳,生成视觉不显的扰动。
  • 多视角对齐优化后,使AI误判率显著提升,身份仍可辨识。
  • 适合数字人、虚拟偶像等需保护隐私的3D人脸应用。

逼真的3D人脸数字资产广泛应用于远程通信、动画和个性化媒体中。然而,视觉语言模型(VLM)无需微调即可通过开放语义推理,从渲染图像中推断敏感属性,带来新隐私风险:一旦3D人脸共享,其任意渲染图都可能被分析出高阶面部特征。现有防御多在2D图像空间操作,无法应对3D表征的身份保持式语义操控。本文提出3D FaceShell,一种在保持几何保真与面部身份一致的前提下,引导VLM对3D模型渲染结果进行语义误导的框架。该方法在原始3D表示上添加可学习的高斯壳,生成空间分布的细微扰动,并通过多视图嵌入对齐进行优化。这些扰动视觉上几乎不可察觉,但能以视图一致的方式有效误导VLM的属性推断。在重建的名人3D人脸数据集及多个黑盒VLM上的实验表明,3D FaceShell显著提升了属性注入和错配率,同时保持了高感知相似性和身份一致性。结果证明,可在不损害人类可识别外观的前提下,操控VLM对3D人脸的语义理解。

原文摘要 · Abstract (English)

Photorealistic 3D face avatars are increasingly deployed as reusable digital assets across applications such as telepresence, animation, and personalized media. At the same time, vision-language models (VLMs) can infer sensitive attributes from rendered images with open-ended semantic reasoning without any fine-tuning. This creates a new privacy challenge: once a 3D face avatar is shared, any of its renderings can be analyzed to extract high-level facial attributes. Existing defenses largely operate in 2D image space and do not address identity-preserving semantic manipulation of 3D facial representations. We propose 3D FaceShell, a framework for steering VLM interpretations of faces rendered from 3D models while preserving geometric fidelity and facial identity. 3D FaceShell augments the original 3D representation with a learnable Gaussian shell that produces subtle, spatially distributed perturbations optimized through multi-view embedding alignment. The perturbations are designed to be visually inconspicuous yet sufficient to redirect VLM-based attribute inference in a view-consistent manner. Extensive experiments on reconstructed celebrity face avatars and multiple black-box VLMs demonstrate that 3D FaceShell significantly increases attribute injection and mismatch rates while maintaining high perceptual similarity and identity consistency. Our results show that it is possible to manipulate VLM-level semantic interpretation of 3D faces without compromising their human-recognizable appearance.

3D人脸隐私保护VLM防御生成对抗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。