通过真值表划分数据,找到二分类器最优线性组合的解析解。
Analytical study of the optimal combination of binary classifiers based on classifiers-induced partitioning of the training set
- 用真值表对数据分组,构建可解析优化的损失函数框架。
- 在三分类器情形下,明确所有存在唯一解、无穷解或无解的配置。
- 直接给出加权系数公式,无需迭代训练,适合高稳定性场景。
本文研究基于真值表对数据集进行逻辑划分后,二分类器的最优线性组合问题。给定的分类器将数据划分为等价类,从而可通过多维推广的分类校准函数,严格分析凸化经验风险。我们建立了任意分类器列表下凸化经验风险全局最小点存在的充分条件与唯一性条件(当分类器数量大时,最小点可能不存在)。在三个分类器的情况下,我们的分析可列出所有导致唯一解、极小值或非唯一最小点的情形。此外,针对指数(Boost)和逻辑(Logit)损失函数,推导出最优权重的显式解析表达式,避免了迭代优化过程。通过引入ϕ-边界概念,可评估所得分类器的稳定性及数据质量。
原文摘要 · Abstract (English)
This paper studies an optimal linear combination of binary classifiers based on a logical structuration of the dataset via truth tables. The given classifiers partition data into equivalence classes, allowing for a rigorous analysis of the convexified empirical risk through a multidimensional generalization of classification calibrated functions. We establish sufficient conditions for the existence and uniqueness of the (global) point of minimum of the convexified empirical risk for any list of classifiers (when the number of classifiers is large, there frequently could be no point of minimum). In the case of three classifiers, our analysis allows to list all the configurations leading to either a unique solution, infima or non-unique points of minimum. Furthermore, we derive explicit analytical formulae for optimal weights using Exponential (Boost) and Logistic (Logit) loss functions, bypassing iterative optimization. The stability of the resulting classifier and the analysis of data quality can be evaluated through the introduction of the notion of $ϕ$-frontiers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。