提出可自适应数据特性的归一化评估指标,更好衡量小样本、不平衡数据上的模型表现。
Developing a Dataset-Adaptive, Normalized Metric for Machine Learning Model Assessment: Integrating Size, Complexity, and Class Imbalance
- 根据数据规模、维度、类别不平衡等特性动态调整评估标准
- 实验验证可准确预测模型在小样本和高维数据中的性能表现
- 适合资源有限场景下的模型选型与优化,尤其对不平衡数据敏感
传统评估指标如准确率、F1分数和精确率在小型、不平衡或高维数据集上可能失效。本文提出一种数据自适应的归一化评估指标,融合数据集大小、特征维度、类别不平衡程度及信噪比等特性。该指标能早期揭示模型在复杂条件下的性能潜力,提供可扩展且灵活的评估框架。通过分类、回归和聚类任务的实验验证,证明其能准确预测模型的可扩展性与性能,确保在数据稀缺场景下的可靠评估。该方法对机器学习工作流中资源分配与模型优化具有重要应用价值。
原文摘要 · Abstract (English)
Traditional metrics like accuracy, F1-score, and precision are frequently used to evaluate machine learning models, however they may not be sufficient for evaluating performance on tiny, unbalanced, or high-dimensional datasets. A dataset-adaptive, normalized metric that incorporates dataset characteristics like size, feature dimensionality, class imbalance, and signal-to-noise ratio is presented in this study. Early insights into the model's performance potential in challenging circumstances are provided by the suggested metric, which offers a scalable and adaptable evaluation framework. The metric's capacity to accurately forecast model scalability and performance is demonstrated via experimental validation spanning classification, regression, and clustering tasks, guaranteeing solid assessments in settings with limited data. This method has important ramifications for effective resource allocation and model optimization in machine learning workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。