提出三种新型多分类损失,理论分析其性能边界。
Structured Proper Loss Geometries for Multiclass Classification: Theory and Controlled Empirical Evaluation
- 设计三类带结构约束的损失函数,优化目标更稳定
- 在标签噪声下表现优于交叉熵,尤其在长尾数据中
- 适合关注鲁棒性与校准性的模型开发者
严格可正评分规则可在总体层面识别真实类别分布,但其曲率会影响优化和有限样本行为。本文研究三种多分类目标:一种类感知二次Bregman评分(CAPM)、一种带受限log-cosh正则项的强凸生成器(HPG),以及一种带退火概率间距惩罚的HPG目标(APMS)。CAPM作为经典二次评分规则的结构化实例,我们推导了其条件遗憾、曲率、取值范围及logit梯度的界;对HPG和APMS分别证明了精确的惩罚范围与条件目标偏移界。通过受控五次种子实验,在Digits、威斯康星乳腺癌数据集及合成的混淆与长尾问题上,评估了干净标签、对称与成对翻转噪声、类别不平衡、校准性、输入噪声及一阶对抗扰动下的表现。候选方法在干净数据上接近交叉熵,在部分噪声场景中表现提升,但五次实验结果为描述性比较,未达显著性证据。在Digits数据集40%对称噪声下,现有噪声标签基线更优;在30:1合成长尾实验中,显式先验调整方法表现更好。消融实验未显示候选方法特有图结构、正则项或间距组件的稳定增益。数学分析验证了理论性质,实验限定了实证证据,二者共同不支持普遍优越性的主张。
原文摘要 · Abstract (English)
Strictly proper scoring rules identify the true conditional class distribution at population level, but their curvature can alter optimization and finite-sample behavior. We study three multiclass objectives: a class-aware quadratic Bregman score (CAPM), a strongly convex generator with constrained log-cosh ridges (HPG), and an HPG objective with an annealed probability-margin penalty (APMS). CAPM is treated as a structured instance of established quadratic scoring-rule theory. We derive conditional-regret, curvature, range, and logit-gradient bounds for CAPM and HPG, and prove exact penalty-range and conditional-target displacement bounds for APMS. Controlled five-seed experiments use Digits, Wisconsin breast cancer, and synthetic confusion and long-tail problems under clean labels, symmetric and pair-flip corruption, class imbalance, calibration evaluation, input corruption, and first-order adversarial perturbations. The candidates are close to cross-entropy on clean data and show descriptive gains in some noisy-label cells, but the five-seed comparisons are interpreted descriptively rather than as significance evidence. The selected noisy-label baselines perform better on Digits with 40% symmetric label noise, and explicit prior-adjustment methods perform better in the 30:1 synthetic long-tail experiment. Ablations do not show a consistent benefit from the candidate-specific graph, ridge, or margin components. The mathematical analysis establishes the stated properties, and the experiments delimit the empirical evidence; together they do not support a claim of general superiority.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。