研究连续数据摘要的多目标对抗攻击与鲁棒防御,提升可信AI上游可靠性。
Toward Trustworthy AI: Multi-Target Adversarial Attacks and Robust Defenses for Continuous Data Summarization

- 基于DR-子模优化构建多目标对抗扰动,破坏摘要代表性。
- 攻击在低至中等预算下显著降低下游任务性能,最高降幅达27.3%。
- 提出带正则化的鲁棒防御策略,适用于真实数据与结构化场景。
可信AI需要可靠的端到端数据处理流程,而不仅是鲁棒的下游预测模型。作为上游组件,数据摘要决定了哪些信息被保留并传递给后续学习或决策模块。因此,对摘要过程的对抗扰动可能以上游方式损害可信AI:改变所选摘要、降低其代表性,并进一步削弱后续学习任务的效用。本文研究在相似性层面扰动下的连续数据摘要对抗攻击,采用DR-子模优化方法。我们证明一类多分辨率图像摘要目标可表示为非负子模集合函数的多元线性扩展,并满足DR-子模性与m-弱单调性。进而将多目标攻击生成建模为一个极小极大问题,其中优化一个可接受的相似性结构扰动以破坏多个目标摘要模型。为缓解此类扰动,我们将针对混合攻击类型的鲁棒防御建模为正则化的极大极小问题。针对两个问题,我们开发了具有理论保证的近似算法。在真实数据与受控聚类基准上的实验表明,所提攻击在代表性低至中等预算范围内有效,可引发下游任务性能损失最高达27.3%。所提防御在结构化设置中提升了鲁棒性与缓解效果的权衡,同时揭示了真实数据上鲁棒保护的参数敏感性。
原文摘要 · Abstract (English)
Trustworthy AI requires reliable data-processing pipelines, not only robust downstream predictive models. As an upstream component, data summarization determines which information is retained and passed to subsequent learning or decision modules. Therefore, adversarial perturbations to the summarization process can compromise trustworthy AI in an upstream manner: they may alter the selected summary, reduce its representativeness, and further degrade the utility of subsequent learning tasks. In this paper, we study adversarial attacks on continuous data summarization under similarity-level perturbations through DR-submodular optimization. We show that a class of multi-resolution image summarization objectives can be formulated as multilinear extensions of non-negative submodular set functions and satisfy DR-submodularity with $m$-weak monotonicity. We then formulate multi-target attack generation as a min-max problem, where one admissible perturbation of the similarity structure is optimized to degrade multiple target summarization models. To mitigate such perturbations, we formulate robust defense against mixed attack types as a regularized max-min problem. For both problems, we develop approximation algorithms with theoretical guarantees. Experiments on real-data and controlled clustered benchmarks show that the proposed attack is effective in representative low-to-moderate budget regimes and can induce downstream task-performance loss. The proposed defense improves the robustness--mitigation trade-off in structured settings, while also revealing the parameter sensitivity of robust protection on real data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。