保护可穿戴多模态系统隐私,防止视觉令牌泄露个人属性。
Defending Wearable VLMs Against Private Attribute Inference

- 在可信设备内对视觉令牌进行残差变换,分离隐私信息。
- 隐私识别准确率从56.7%降至7.4%,实用度仍保持74.4%。
- 适合关注可穿戴AI隐私安全的研究者与开发者。
可穿戴视觉-语言模型(VLM)通过第一人称视觉捕捉提供持续的多模态辅助:用户针对周围场景提问,系统利用紧凑的视觉令牌支持语言推理。然而,用于辅助的同一视觉证据也可能暴露佩戴者或附近他人的私人属性。本文将此问题定义为分层推理中的隐私-效用权衡,其中视觉编码在可信设备边界内完成,但中间视觉令牌可能传至下游推理组件。这暴露了一个被忽视的泄露面:即使最终文本响应无害,外部攻击者或不受信的下游组件仍可从传输的视觉令牌中恢复私人属性。为此,我们构建了一个包含3,221条图像-问题记录的配对隐私-效用基准,每条均配有用途问题和涵盖位置、收入、性别、兴趣的隐私标签。进一步提出一种预大模型的令牌解耦方法TGAP,学习视觉令牌在离开可信边界前的残差变换。TGAP结合效用保留、身份正则化、语义隐私抑制和图像驱动表征抑制,避免了粗粒度硬掩码或注意力掩码带来的效用损失。在源模型评估基准上,TGAP将隐私识别准确率从56.7%降至7.4%,绝对下降49.3%,同时维持宽松效用达74.4%。结果表明,保护紧凑令牌接口是实现可穿戴多模态AI隐私安全的一条可行路径。
原文摘要 · Abstract (English)
Wearable VLM pipelines promise continuous multimodal assistance from egocentric visual capture: a user asks a task-driven question about the surrounding scene, and the system uses compact visual tokens to support language reasoning. The challenge motivating this work is that the same egocentric evidence needed for useful assistance can also reveal private attributes about the wearer or nearby bystanders. We investigate this as a joint privacy-utility problem for split VLM inference, where visual encoding occurs within a trusted device boundary but intermediate visual tokens may be transmitted to downstream reasoning components. This exposes an understudied leakage surface: even when final textual responses are benign, external attackers or untrusted downstream components can recover private attributes from transmitted visual tokens. To evaluate this tension, we construct a paired privacy-utility benchmark with 3,221 image-question records, each paired with a utility question and privacy labels covering location, income, sex, and interests. We further propose Token-Guided Attribute Privacy (TGAP), a pre-LLM token disentangler that learns a residual transformation of visual tokens before they leave the trusted boundary. TGAP combines utility preservation, identity regularization, semantic privacy suppression, and image-driven representation suppression, avoiding the utility loss caused by coarse hard or attention masking. On the benchmark used for source-model evaluation, TGAP reduces privacy accuracy from 56.7\% to 7.4\%, a 49.3\% absolute drop, while maintaining relaxed utility at 74.4\%. These results suggest that securing the compact token interface is a practical path toward privacy-preserving wearable multimodal AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。