用视觉语言模型识别增强现实中的隐私风险,自动隐藏敏感信息。
See No Evil: Semantic Context-Aware Privacy Risk Detection for AR

- 通过链式思维提示让模型理解场景语义,识别如密码便签等敏感内容。
- 在真实AR数据集上达到81.48%准确率和84.62%F1分数,隐私泄露率降至17.58%。
- 设计智能提醒界面,帮助用户提升对隐私风险的认知。
增强现实(AR)系统因持续捕获视觉数据而带来独特隐私风险。现有框架缺乏对视觉内容的语义理解,难以有效识别上下文相关的隐私威胁。本文提出PrivAR,利用视觉语言模型(VLMs)结合链式思维提示,实现AR环境中的情境化隐私风险检测。PrivAR通过视觉场景线索推断潜在敏感信息类型,例如在办公环境中识别密码便签。该方法可检测并模糊文本内容,防止敏感信息暴露,同时保留支撑VLM推理所需的上下文线索。此外,我们研究了情境感知的警告界面,以提升用户隐私意识。在真实世界AR数据集上的实验表明,PrivAR相较基线方法取得更高准确率(81.48%)与F1分数(84.62%),隐私泄露率降低至17.58%。用户研究进一步揭示了有效隐私感知AR设计的关键因素。
原文摘要 · Abstract (English)
Augmented reality (AR) systems pose unique privacy risks due to their continuous capture of visual data. Existing AR privacy frameworks lack semantic understanding of visual content, limiting their effectiveness in detecting context-dependent privacy risks. We propose PrivAR, which leverages vision language models (VLMs) with chain-of-thought prompting for contextual privacy risk detection in AR environments. PrivAR uses visual scene cues to infer potential sensitive information types, such as identifying password notes in office environments through contextual reasoning. PrivAR detects and obfuscates textual content, preventing exposure of sensitive information while preserving contextual cues necessary for VLM inference. Additionally, we investigate contextually-informed warning interfaces to enhance user privacy awareness. Experiments on a real-world AR dataset show that PrivAR achieves superior accuracy (81.48%) and F1-score (84.62%) compared to baselines, while reducing privacy leakage rate to 17.58%. User studies evaluating contextually-informed warning interfaces provide insights into effective privacy-aware AR design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。