根据用户偏好动态调整评估指标,精准识别模型在不同领域的表现瓶颈。
Multi-domain performance analysis with scores tailored to user preferences
- 用概率框架将性能建模为分布,基于用户偏好加权平均
- 发现排名类评分能保证加权平均结果与整体性能一致
- 提出易、难、主导、瓶颈四类领域,适配个性化评估需求
算法性能高度依赖其应用领域的数据分布,而不同领域间的分布特性各异。在多领域评估后,计算加权平均性能虽常见,但深入分析该平均过程更具价值。本文采用概率框架,将性能视为概率测度(如分类任务的归一化混淆矩阵),指出加权平均实为性能的汇总。只有特定评分方法——即参数化于用户偏好的排名类分数——能使汇总性能等于各领域性能的加权算术平均。据此,本文严格定义了四类领域:最易、最难、主导和瓶颈领域,其划分完全依赖用户偏好。理论构建不依赖具体任务,随后针对二分类问题开发了新型可视化工具。
原文摘要 · Abstract (English)
The performance of algorithms, methods, and models tends to depend heavily on the distribution of cases on which they are applied, this distribution being specific to the applicative domain. After performing an evaluation in several domains, it is highly informative to compute a (weighted) mean performance and, as shown in this paper, to scrutinize what happens during this averaging. To achieve this goal, we adopt a probabilistic framework and consider a performance as a probability measure (e.g., a normalized confusion matrix for a classification task). It appears that the corresponding weighted mean is known to be the summarization, and that only some remarkable scores assign to the summarized performance a value equal to a weighted arithmetic mean of the values assigned to the domain-specific performances. These scores include the family of ranking scores, a continuum parameterized by user preferences, and that the weights to consider in the arithmetic mean depend on the user preferences. Based on this, we rigorously define four domains, named easiest, most difficult, preponderant, and bottleneck domains, as functions of user preferences. After establishing the theory in a general setting, regardless of the task, we develop new visual tools for two-class classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。