arXiv:2511.11767cs.LGcs.CY2025-11AAAI

用可解释的KAN网络提升模型公平性,兼顾准确率与透明度。

Learning Fair Representations with Kolmogorov-Arnold Networks

  • 基于样条的KAN网络增强对抗学习稳定性
  • 自适应公平惩罚机制实现公平与精度平衡
  • 适合高风险决策场景的可解释公平建模

尽管公平感知机器学习取得进展,预测模型仍常对边缘群体表现出歧视行为。这种不公平可能源于训练数据偏差、模型设计或群体间表征差异,在大学招生等高风险领域带来挑战。现有公平学习模型难以在公平性与准确性间取得最优平衡,且依赖黑箱模型限制了可解释性。为此,本文将科尔莫戈罗夫-阿诺德网络(KANs)融入公平对抗学习框架,利用其对抗鲁棒性与可解释性,实现稳定的对抗学习。我们推导了基于样条的KAN架构在对抗优化中的理论稳定性,并提出自适应公平惩罚更新机制,以平衡公平性与准确性。在两个真实招生数据集上的实证结果表明,该框架能在保持预测性能的同时,有效提升对敏感属性的公平性表现。

原文摘要 · Abstract (English)

Despite recent advances in fairness-aware machine learning, predictive models often exhibit discriminatory behavior towards marginalized groups. Such unfairness might arise from biased training data, model design, or representational disparities across groups, posing significant challenges in high-stakes decision-making domains such as college admissions. While existing fair learning models aim to mitigate bias, achieving an optimal trade-off between fairness and accuracy remains a challenge. Moreover, the reliance on black-box models hinders interpretability, limiting their applicability in socially sensitive domains. To circumvent these issues, we propose integrating Kolmogorov-Arnold Networks (KANs) within a fair adversarial learning framework. Leveraging the adversarial robustness and interpretability of KANs, our approach facilitates stable adversarial learning. We derive theoretical insights into the spline-based KAN architecture that ensure stability during adversarial optimization. Additionally, an adaptive fairness penalty update mechanism is proposed to strike a balance between fairness and accuracy. We back these findings with empirical evidence on two real-world admissions datasets, demonstrating the proposed framework's efficiency in achieving fairness across sensitive attributes while preserving predictive performance.

公平学习可解释性对抗学习KAN

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。