arXiv:2604.11704cs.LGcs.AI2026-04

通过几何方法识别并清除模型捷径,提升公平性与泛化能力。

Fairness is Not Flat: Geometric Phase Transitions Against Shortcut Learning

  • 用无隐藏层的拓扑审计器自动识别主导梯度的冗余特征。
  • 剪枝线性捷径后,模型需使用更高几何容量(N≥16)学习更复杂决策边界。
  • 相比传统方法,计算成本低,性别偏差从21.18%降至7.66%。

深度神经网络极易陷入捷径学习,常记忆低维伪相关而非底层因果机制,不仅降低分布外鲁棒性,还在敏感应用中引发严重人口偏差。本文提出一种几何先验方法以缓解此问题。通过部署零隐藏层(N=1)的拓扑审计器,数学上隔离不依赖人工干预的主导梯度特征。实证发现容量相变现象:一旦线性捷径被剪枝,网络被迫利用更高几何容量(N≥16)弯曲决策边界,学习更符合伦理的表示。该方法优于L1正则化(会引发人口偏差),且计算成本仅为事后方法如Just Train Twice(JTT)的一小部分,成功将反事实性别脆弱性从21.18%降至7.66%。

原文摘要 · Abstract (English)

Deep Neural Networks are highly susceptible to shortcut learning, frequently memorizing low-dimensional spurious correlations instead of underlying causal mechanisms. This phenomenon not only degrades out-of-distribution robustness but also induces severe demographic biases in sensitive applications. In this paper, we propose a geometric \textit{a priori} methodology to mitigate shortcut learning. By deploying a zero-hidden-layer ($N=1$) Topological Auditor, we mathematically isolate features that monopolize the gradient without human intervention. We empirically demonstrate a Capacity Phase Transition: once linear shortcuts are pruned, networks are forced to utilize higher geometric capacity ($N \geq 16$) to curve the decision boundary and learn ethical representations. Our approach outperforms L1 Regularization -- which collapses into demographic bias -- and operates at a fraction of the computational cost of post-hoc methods like Just Train Twice (JTT), successfully reducing counterfactual gender vulnerability from 21.18\% to 7.66\%.

公平性捷径学习几何容量拓扑审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。