为分类模型定义可信赖的安全决策区域,确保预测可靠性。
Exact characterization of ε-Safe Decision Regions for exponential family distributions and Multi Cost SVM approximation
- 提出ε-安全区域定义,基于指数族分布可精确计算
- 在指数族数据下,安全区域可通过参数设计控制
- 开发多代价SVM近似方法,适用于不平衡数据
数据驱动分类器的预测需具备概率保证,以定义可信赖模型。本文首次正式定义ε-安全决策区域——输入空间中对目标类别预测具有概率保障的子集。当数据来自指数族分布时,该区域的形式可被解析确定,并通过目标类别采样概率与预测置信度等设计参数进行调控。然而,指数族假设并非总成立。为此,本文提出多代价支持向量机(Multi Cost SVM),一种基于SVM的算法,可近似安全区域并处理数据不平衡问题。研究包含实验验证与代码开源,确保可复现性。
原文摘要 · Abstract (English)
Probabilistic guarantees on the prediction of data-driven classifiers are necessary to define models that can be considered reliable. This is a key requirement for modern machine learning in which the goodness of a system is measured in terms of trustworthiness, clearly dividing what is safe from what is unsafe. The spirit of this paper is exactly in this direction. First, we introduce a formal definition of ε-Safe Decision Region, a subset of the input space in which the prediction of a target (safe) class is probabilistically guaranteed. Second, we prove that, when data come from exponential family distributions, the form of such a region is analytically determined and controllable by design parameters, i.e. the probability of sampling the target class and the confidence on the prediction. However, the request of having exponential data is not always possible. Inspired by this limitation, we developed Multi Cost SVM, an SVM based algorithm that approximates the safe region and is also able to handle unbalanced data. The research is complemented by experiments and code available for reproducibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。