构建文档政策保护基准,揭示大模型推理中敏感信息泄露问题
Doc-PP: Document Policy Preservation Benchmark for Large Vision-Language Models
- 提出DVA框架,分离推理与政策验证过程
- 实测显示复杂跨模态推理导致47%敏感信息泄露
- 适合关注文档安全与多模态合规的开发者
大型视觉语言模型(LVLMs)在真实文档问答中的部署常受限于动态用户定义的披露政策。现有安全研究多聚焦隐含社会规范或纯文本场景,忽视了多模态文档的复杂性。本文提出Doc-PP(文档政策保留基准),基于真实报告构建,需在严格非披露政策下对异构视觉与文本元素进行推理。评估发现存在系统性推理诱发的安全缺口:当答案需通过复杂合成或跨模态聚合推断时,模型频繁泄露敏感信息,绕过现有安全约束。此外,提取文本虽提升感知能力,却意外助长泄露。为此,我们提出DVA(分解-验证-聚合)结构化推理框架,将推理与策略验证解耦。实验表明,DVA显著优于标准提示防御方法,为政策合规文档理解提供稳健基线。
原文摘要 · Abstract (English)
The deployment of Large Vision-Language Models (LVLMs) for real-world document question answering is often constrained by dynamic, user-defined policies that dictate information disclosure based on context. While ensuring adherence to these explicit constraints is critical, existing safety research primarily focuses on implicit social norms or text-only settings, overlooking the complexities of multimodal documents. In this paper, we introduce Doc-PP (Document Policy Preservation Benchmark), a novel benchmark constructed from real-world reports requiring reasoning across heterogeneous visual and textual elements under strict non-disclosure policies. Our evaluation highlights a systemic Reasoning-Induced Safety Gap: models frequently leak sensitive information when answers must be inferred through complex synthesis or aggregated across modalities, effectively circumventing existing safety constraints. Furthermore, we identify that providing extracted text improves perception but inadvertently facilitates leakage. To address these vulnerabilities, we propose DVA (Decompose-Verify-Aggregation), a structural inference framework that decouples reasoning from policy verification. Experimental results demonstrate that DVA significantly outperforms standard prompting defenses, offering a robust baseline for policy-compliant document understanding
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。