arXiv:2601.04941cs.LG2026-01

用数学中的基数概念改进损失函数,缓解数据不均衡问题。

Cardinality augmented loss functions

  • 基于度量空间的有效多样性设计新型损失函数
  • 在人工和真实材料科学数据上显著提升少数类表现
  • 适合处理类别不平衡的机器学习任务

类别不平衡是神经网络训练中的常见且严重问题,多数类常主导训练过程,导致分类器性能偏向多数类别。为此,本文引入基数增强型损失函数,其灵感来自现代数学中诸如“度量”和“扩散”等基数相关不变量,这些不变量通过评估度量空间的‘有效多样性’来扩展基数概念,因而成为解决训练数据过于同质化的自然方案。本文建立了将基数增强损失函数应用于神经网络训练的方法,并在人工构造的不平衡数据集及真实世界材料科学不平衡数据集上进行了实验。结果表明,少数类的性能显著提升,整体性能指标也得到改善。

原文摘要 · Abstract (English)

Class imbalance is a common and pernicious issue for the training of neural networks. Often, an imbalanced majority class can dominate training to skew classifier performance towards the majority outcome. To address this problem we introduce cardinality augmented loss functions, derived from cardinality-like invariants in modern mathematics literature such as magnitude and the spread. These invariants enrich the concept of cardinality by evaluating the `effective diversity' of a metric space, and as such represent a natural solution to overly homogeneous training data. In this work, we establish a methodology for applying cardinality augmented loss functions in the training of neural networks and report results on both artificially imbalanced datasets as well as a real-world imbalanced material science dataset. We observe significant performance improvement among minority classes, as well as improvement in overall performance metrics.

损失函数类别不平衡度量学习神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。