首个支持图文混合输入的鲁棒性认证框架,能同时应对图像和文本扰动。
Certified Robustness under Heterogeneous Perturbations via Hybrid Randomized Smoothing

- 基于奈曼-皮尔逊准则构建联合最坏情况分析,统一处理离散与连续输入。
- 推出闭式一维认证结果,比单一模态方法更严格且通用。
- 适用于交互依赖的图文安全过滤,为多模态模型提供可信赖保障。
随机平滑能提供强而通用的鲁棒性证明,但现有方法仅限于单一模态,对连续与离散输入分别处理。在多模态模型中,决策依赖跨模态语义,攻击者可联合扰动异构输入,导致单模态认证失效。本文提出一种统一的随机平滑框架,用于混合离散-连续输入,基于可解析的奈曼-皮尔逊联合最坏情况公式。通过分析因子化噪声诱导的联合似然排序,该方法获得闭式的一维认证结果,严格推广了高斯(图像仅)和离散(文本仅)随机平滑。我们在多模态安全过滤任务上验证了该框架,首次实现了模型无关的奈曼-皮尔逊认证,覆盖交互依赖的文本-图像联合扰动。
原文摘要 · Abstract (English)
Randomized smoothing provides strong, model-agnostic robustness certificates, but existing guarantees are limited to single modalities, treating continuous and discrete inputs in isolation. This limitation becomes critical in multimodal models, where decisions depend on cross-modal semantics and adversaries can jointly perturb heterogeneous inputs, rendering unimodal certificates insufficient. We introduce a unified randomized smoothing framework for mixed discrete--continuous inputs based on an analytically tractable Neyman--Pearson formulation of the joint worst-case problem. By analyzing the joint likelihood ordering induced by factorized discrete and continuous noise, our approach yields a closed-form, one-dimensional certificate that strictly generalizes both Gaussian (image-only) and discrete (text-only) randomized smoothing. We validate the framework on multimodal safety filtering, providing, to our knowledge, the first model-agnostic Neyman--Pearson certificate for joint discrete-token and continuous-image perturbations in interaction-dependent text--image safety filtering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。