arXiv:2606.19212stat.MLcs.LG2026-06被引 1

揭示语义扰动下模型误判的几何本质,给出可计算的攻击风险指标。

Generalised Eigenvalue Geometry of Semantic Adversarial Attacks

论文配图:Generalised Eigenvalue Geometry of Semantic Adversarial Attacks
图 1 · 摘自论文原文
  • 基于雅可比矩阵构建广义特征值模型,刻画双模型语义扰动机制。
  • 最大广义特征值决定局部表征偏移上限,导出预测翻转的闭式条件。
  • 适用于金融文本分类器的攻击风险评估与离散搜索验证框架。

近期实证研究表明,语义等价的改写可能使金融情感分类器失效:尽管改写文本在强参考嵌入下与原句相近,却足以改变目标模型的表征并导致分类错误。现有鲁棒性理论或假设单一模型威胁模型,或仅关注经验攻击算法。本文建立连续局部语义改写扰动模型,捕捉双模型结构。我们证明,在代理模型预算约束下,目标表征最坏局部偏移由矩阵束 (A,B) 的最大广义特征值决定,其中 A、B 来自两个嵌入映射的雅可比矩阵。由此得到的攻击性指数 λ*(x) 是局部改写几何与嵌入器选择的内在属性,可对仿射读出层给出闭式预测翻转条件,并支持保守的群体与有限样本攻击性证书。为统一控制仿射读出类别的攻击性,我们推导出二元攻击性指标的无分布 VC 界,以及基于攻击性调整边界(扣除局部几何惩罚)的尺度敏感边界。我们还将连续理论与离散改写搜索相联系,识别成功与失败有限搜索间的不对称性,并给出离散与连续设置一致的覆盖条件。最后,提出基于软词元松弛与生成改写集的实证验证框架,用于评估部署中的金融文本分类器的局部特征值几何、预测翻转条件及有限搜索近似效果。

原文摘要 · Abstract (English)

Recent empirical work shows that semantically equivalent paraphrases can fool financial sentiment classifiers: although a paraphrase remains close to the original under a strong reference embedding, it may shift the target model's representation enough to change the predicted class. Existing robustness theory either assumes a single-model threat model or focuses mainly on empirical attack algorithms. We develop a continuous local model of semantic paraphrase perturbations that captures this two-model structure. We show that the worst-case local displacement of the target representation, subject to a proxy-model budget, is governed by the largest generalised eigenvalue of a matrix pencil $(A,B)$ constructed from the Jacobians of the two embedding maps. The resulting attackability index $λ^*(x)$ is intrinsic to the local paraphrase geometry and the chosen embedders, yields a closed-form prediction-flip condition for affine readouts, and supports conservative population and finite-sample attackability certificates. For uniform control over classes of affine readouts, we derive a distribution-free VC bound for binary attackability indicators and a scale-sensitive margin bound based on an attackability-adjusted margin that subtracts a local geometric penalty from the standard classifier margin. We also connect the continuous theory to discrete paraphrase search, identify an asymmetry between successful and unsuccessful finite searches, and give a covering condition under which the discrete and continuous settings agree. Finally, we propose an empirical verification framework using soft-token relaxations and generated paraphrase sets to assess the local eigenvalue geometry, prediction-flip condition, and finite-search approximation on a deployed financial-text classifier.

对抗攻击语义扰动金融文本特征值分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。