arXiv:2507.03392cs.LG2025-07综述被引 5

梳理机器学习绝对评估指标,帮助跨模型跨任务精准比较。

Absolute Evaluation Measures for Machine Learning: A Survey

  • 按学习任务分类整理绝对评估方法,统一评价尺度。
  • 涵盖分类、聚类、回归、排序四类任务的评估指标。
  • 适合需要公平比较模型性能的研究者与工程师。

机器学习广泛应用于计算机科学、社会科学、医学、化学和金融等多个领域,其多样性导致评估方法各异,难以有效比较模型性能。绝对评估指标通过在固定尺度上衡量模型表现,不依赖参考模型和数据范围,实现明确比较。然而,许多常用指标不具备普适性,缺乏使用指导。本文综述了机器学习中的绝对评估指标,按学习任务类型组织,覆盖分类、聚类、回归和排序四类。通过针对不同任务挑战分类指标,为实践者提供选型依据,提升个体模型评估质量,并促进跨模型、跨应用的有意义比较。

原文摘要 · Abstract (English)

Machine Learning is a diverse field applied across various domains such as computer science, social sciences, medicine, chemistry, and finance. This diversity results in varied evaluation approaches, making it difficult to compare models effectively. Absolute evaluation measures offer a practical solution by assessing a model's performance on a fixed scale, independent of reference models and data ranges, enabling explicit comparisons. However, many commonly used measures are not universally applicable, leading to a lack of comprehensive guidance on their appropriate use. This survey addresses this gap by providing an overview of absolute evaluation metrics in ML, organized by the type of learning problem. While classification metrics have been extensively studied, this work also covers clustering, regression, and ranking metrics. By grouping these measures according to the specific ML challenges they address, this survey aims to equip practitioners with the tools necessary to select appropriate metrics for their models. The provided overview thus improves individual model evaluation and facilitates meaningful comparisons across different models and applications.

评估指标机器学习综述模型比较

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。