剖析Softmax家族损失函数,揭示其分类与排序一致性及收敛特性。
Is Softmax Loss All You Need? A Principled Analysis of Softmax-family Loss
- 基于Fenchel-Young框架系统分析Softmax族损失的理论性质。
- 发现不同损失在收敛行为上存在显著差异,影响模型性能。
- 提供可解释的偏差-方差分解,指导大规模分类中的损失选择。
Softmax损失是分类与排序任务中最广泛使用的代理目标之一。为阐明其理论特性,Fenchel-Young框架将其置于一个广义代理损失族的典型位置。与此同时,另一研究方向聚焦于类别数量极多时的可扩展性问题,提出了多种近似方法,在保留精确目标优势的同时提升效率。本文结合这两条路径,对Softmax族损失进行系统性分析:检验不同代理损失是否与分类和排序指标一致,并通过梯度动态分析揭示其不同的收敛行为。我们还提出一种针对近似方法的系统性偏差-方差分解,提供收敛保证,并进一步推导出每轮迭代的复杂度分析,明确揭示有效性和效率之间的权衡。在代表性任务上的大量实验表明,一致性、收敛性与实际性能之间存在强关联。这些结果共同构建了大类别机器学习中损失选择的理论基础,并提供实用指导。
原文摘要 · Abstract (English)
The Softmax loss is one of the most widely employed surrogate objectives for classification and ranking tasks. To elucidate its theoretical properties, the Fenchel-Young framework situates it as a canonical instance within a broad family of surrogates. Concurrently, another line of research has addressed scalability when the number of classes is exceedingly large, in which numerous approximations have been proposed to retain the benefits of the exact objective while improving efficiency. Building on these two perspectives, we present a principled investigation of the Softmax-family losses. We examine whether different surrogates achieve consistency with classification and ranking metrics, and analyze their gradient dynamics to reveal distinct convergence behaviors. We also introduce a systematic bias-variance decomposition for approximate methods that provides convergence guarantees, and further derive a per-epoch complexity analysis, showing explicit trade-offs between effectiveness and efficiency. Extensive experiments on a representative task demonstrate a strong alignment between consistency, convergence, and empirical performance. Together, these results establish a principled foundation and offer practical guidance for loss selections in large-class machine learning applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。