arXiv:2602.18154cs.CLcs.AI2026-02中稿 · paper

构建首个金融领域多模态越狱检测数据集,助力安全AI落地。

FENCE: A Financial and Multimodal Jailbreak Detection Dataset

  • 聚焦金融场景,融合文本与图像构造真实越狱攻击
  • 测试显示GPT-4o存在可测量攻击成功率,开源模型更易受攻
  • 提供中英双语数据,适配金融AI安全研发与评估

越狱攻击对大型语言模型(LLMs)和视觉语言模型(VLMs)的部署构成重大风险。由于同时处理文本与图像,VLMs 更易受攻击。然而,现有越狱检测资源匮乏,尤其在金融领域。为此,我们提出 FENCE——一个中英文双语、面向金融应用的多模态越狱检测数据集,用于训练与评估检测器。FENCE 通过与金融相关查询结合图像驱动威胁,强调领域真实性。对商用与开源 VLM 的实验表明其普遍存在漏洞:GPT-4o 出现可测量的攻击成功,开源模型暴露程度更高。基于 FENCE 训练的基准检测器在分布内测试中达到 99% 准确率,并在外部基准上保持强性能,证明其训练可靠性。FENCE 为金融领域多模态越狱检测提供了精准资源,支持更安全、可靠的 AI 系统建设。警告:本文包含可能令人不适的示例数据。

原文摘要 · Abstract (English)

Jailbreaking poses a significant risk to the deployment of Large Language Models (LLMs) and Vision Language Models (VLMs). VLMs are particularly vulnerable because they process both text and images, creating broader attack surfaces. However, available resources for jailbreak detection are scarce, particularly in finance. To address this gap, we present FENCE, a bilingual (Korean-English) multimodal dataset for training and evaluating jailbreak detectors in financial applications. FENCE emphasizes domain realism through finance-relevant queries paired with image-grounded threats. Experiments with commercial and open-source VLMs reveal consistent vulnerabilities, with GPT-4o showing measurable attack success rates and open-source models displaying greater exposure. A baseline detector trained on FENCE achieves 99 percent in-distribution accuracy and maintains strong performance on external benchmarks, underscoring the dataset's robustness for training reliable detection models. FENCE provides a focused resource for advancing multimodal jailbreak detection in finance and for supporting safer, more reliable AI systems in sensitive domains. Warning: This paper includes example data that may be offensive.

越狱检测金融AI多模态安全评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。