给视觉语言模型的隐私理解能力建了新评估体系
Assessing Visual Privacy Risks in Multimodal AI: A Novel Taxonomy-Grounded Evaluation of Vision-Language Models
- 构建多层级视觉隐私分类体系,覆盖多种隐私风险场景
- 测试多个主流视觉语言模型,发现其隐私理解能力参差不齐
- 为未来隐私感知AI研发提供基准和方向,适合安全与伦理研究者
近年来,人工智能深刻改变了技术格局。大型语言模型在推理、文本理解、上下文模式识别以及图文融合理解方面表现出色。然而,这些进展暴露出模型对隐私概念认知的严重不足。因此,迫切需要明确这些模型是否以及如何理解并执行隐私原则,尤其在缺乏相关测试资源的情况下。本文通过探讨法律框架如何指导新兴技术的能力,提出一个全面、多层次的视觉隐私分类体系,可扩展且适应现有及未来研究需求。同时,我们评估了多个前沿视觉语言模型,揭示其在情境隐私理解上存在显著不一致。本工作既为未来研究奠定了分类基础,也提供了当前模型局限性的关键基准,凸显了开发更强大、具备隐私意识的AI系统的紧迫性。
原文摘要 · Abstract (English)
Artificial Intelligence have profoundly transformed the technological landscape in recent years. Large Language Models (LLMs) have demonstrated impressive abilities in reasoning, text comprehension, contextual pattern recognition, and integrating language with visual understanding. While these advances offer significant benefits, they also reveal critical limitations in the models' ability to grasp the notion of privacy. There is hence substantial interest in determining if and how these models can understand and enforce privacy principles, particularly given the lack of supporting resources to test such a task. In this work, we address these challenges by examining how legal frameworks can inform the capabilities of these emerging technologies. To this end, we introduce a comprehensive, multi-level Visual Privacy Taxonomy that captures a wide range of privacy issues, designed to be scalable and adaptable to existing and future research needs. Furthermore, we evaluate the capabilities of several state-of-the-art Vision-Language Models (VLMs), revealing significant inconsistencies in their understanding of contextual privacy. Our work contributes both a foundational taxonomy for future research and a critical benchmark of current model limitations, demonstrating the urgent need for more robust, privacy-aware AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。