arXiv:2601.08235cs.AIcs.CL2026-01被引 2

首个评估智能体多模态隐私行为的基准,揭示模型在保护隐私与实用间的关键失衡。

MPCI-Bench: A Benchmark for Multimodal Pairwise Contextual Integrity Evaluation of Language Model Agents

  • 构建多模态成对场景,涵盖规范判断、故事推理和真实动作轨迹三层次评估。
  • 实测主流模型在视觉隐私泄露上比文本更严重,且难以平衡隐私与功能需求。
  • 适合研究智能体隐私安全、多模态生成与可信AI的学者与开发者参考。

随着语言模型代理从被动聊天机器人演变为处理个人数据的主动助手,评估其遵守社会规范的能力变得日益重要,常以情境完整性(Contextual Integrity, CI)为视角。然而,现有CI基准大多局限于文本,侧重负面拒绝场景,忽视了多模态隐私风险以及隐私与效用之间的根本权衡。本文提出MPCI-Bench,首个用于评估代理场景下多模态情境完整性的基准。该基准包含源自同一视觉源的成对正负样本,覆盖三个层级:规范性种子判断、富含上下文的故事推理,以及可执行的代理行为轨迹。通过三原则迭代精炼流程保障数据质量。对先进多模态模型的评估发现,它们普遍存在隐私与效用失衡问题,且存在显著的模态泄漏差距——敏感视觉信息泄露频率高于文本信息。我们将开源MPCI-Bench,以推动代理情境完整性研究的发展。

原文摘要 · Abstract (English)

As language-model agents evolve from passive chatbots into proactive assistants that handle personal data, evaluating their adherence to social norms becomes increasingly critical, often through the lens of Contextual Integrity (CI). However, existing CI benchmarks are largely text-centric and primarily emphasize negative refusal scenarios, overlooking multimodal privacy risks and the fundamental trade-off between privacy and utility. In this paper, we introduce MPCI-Bench, the first Multimodal Pairwise Contextual Integrity benchmark for evaluating privacy behavior in agentic settings. MPCI-Bench consists of paired positive and negative instances derived from the same visual source and instantiated across three tiers: normative Seed judgments, context-rich Story reasoning, and executable agent action Traces. Data quality is ensured through a Tri-Principle Iterative Refinement pipeline. Evaluations of state-of-the-art multimodal models reveal systematic failures to balance privacy and utility and a pronounced modality leakage gap, where sensitive visual information is leaked more frequently than textual information. We will open-source MPCI-Bench to facilitate future research on agentic CI.

多模态隐私评估智能体情境完整性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。