测试12个视觉语言模型在真实物理场景中的隐私感知能力,发现普遍存在识别缺陷。
How Far Are VLMs from Privacy Awareness in the Physical World? An Empirical Study

- 构建音视频交互框架ImmersionPrivacy,模拟家庭医院等真实环境
- 复杂场景下模型准确率随混乱度上升而持续下降,社交情境切换后最高仅65%正确率
- 面对任务指令与隐私约束冲突时,最优模型也仅在51%情况下平衡二者
随着视觉语言模型(VLMs)被广泛部署为具身助手的认知核心,评估其在物理环境中的隐私感知能力变得至关重要。与数字聊天机器人不同,这些智能体在家庭、医院等私密空间中具备物理观察和操作能力,可能接触敏感信息。然而现有评测仍局限于单模态文本,无法反映真实场景需求。为此,我们提出ImmersionPrivacy——基于Unity的交互式音视频评估框架,通过三阶段渐进测试模型在杂乱场景中识别敏感物品、适应社会情境变化、处理显性指令与隐含隐私约束冲突的能力。对12个先进模型的评估显示:在复杂场景中,所有模型性能随场景混乱度增加而单调下降;社会情境变化时,无一模型准确率超过65%;面对冲突指令,最佳模型gemini-3.1-pro仅在51%案例中成功兼顾任务完成与隐私保护。结果表明,当前VLM在物理世界中存在感知脆弱性,其隐私认知无法有效指导实际行为。代码与数据已开源。
原文摘要 · Abstract (English)
As Vision-Language Models (VLMs) are increasingly deployed as autonomous cognitive cores for embodied assistants, evaluating their privacy awareness in physical environments becomes critical. Unlike digital chatbots, these agents operate in intimate spaces, such as homes and hospitals, where they possess the physical agency to observe and manipulate privacy-sensitive information and artifacts. However, current benchmarks remain limited to unimodal, text-based representations that cannot capture the demands of real-world settings. To bridge this gap, we present ImmersedPrivacy, an interactive audio-visual evaluation framework that simulates realistic physical environments using a Unity-based simulator. ImmersedPrivacy evaluates physically grounded privacy awareness across three progressive tiers that test a model's ability to identify sensitive items in cluttered scenes, adapt to shifting social contexts, and resolve conflicts between explicit commands and inferred privacy constraints. Our evaluation of 12 state-of-the-art models reveals consistent deficits. In cluttered scenes, all models exhibit monotonic performance decay as scene complexity grows due to perceptual deficit. When social context shifts, no model exceed 65% selection accuracy. Under conflicting commands, the best model gemini-3.1-pro perfectly balances task completion and privacy preservation in only 51% of cases. These findings reveal that current VLMs in the physical world suffer from perceptual fragility and fail to let their knowledge of privacy cues govern their situated behavior. Our code and data is available at https://github.com/immersed-privacy/immersed-privacy .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。