arXiv:2501.15977cs.LGstat.ML2025-01中稿 · ICASSP 2025被引 1

分析低贝叶斯误差下分类误差与模型差异的关系,给出可解释的误差上界。

Classification Error Bound for Low Bayes Error Conditions in Machine Learning

  • 在低贝叶斯误差条件下,提出分类误差上界的线性近似方法。
  • 揭示了分类误差与KL散度之间的定量关系,适用于真实分布未知场景。
  • 适用于语音识别等任务,帮助理解误差率、交叉熵和困惑度的关联。

在统计分类与机器学习中,分类误差是重要性能指标,其最小值由贝叶斯决策规则决定。实际中,真实分布通常被训练数据估计的模型分布替代,导致贝叶斯误差与基于模型的分类误差之间存在偏差。本文利用分类误差上界研究该偏差与KL散度的关系。针对近期许多机器学习任务中模型分类误差较低的现象,假设贝叶斯误差较小,提出低贝叶斯误差条件下的分类误差上界线性近似。进一步讨论了先验概率对边界的影响,并将误差上界扩展至序列任务。以自动语音识别为例,从理论上分析了交叉熵损失、语言模型困惑度与词错误率之间的关联,通过扩展的上界实现多指标统一建模。

原文摘要 · Abstract (English)

In statistical classification and machine learning, classification error is an important performance measure, which is minimized by the Bayes decision rule. In practice, the unknown true distribution is usually replaced with a model distribution estimated from the training data in the Bayes decision rule. This substitution introduces a mismatch between the Bayes error and the model-based classification error. In this work, we apply classification error bounds to study the relationship between the error mismatch and the Kullback-Leibler divergence in machine learning. Motivated by recent observations of low model-based classification errors in many machine learning tasks, bounding the Bayes error to be lower, we propose a linear approximation of the classification error bound for low Bayes error conditions. Then, the bound for class priors are discussed. Moreover, we extend the classification error bound for sequences. Using automatic speech recognition as a representative example of machine learning applications, this work analytically discusses the correlations among different performance measures with extended bounds, including cross-entropy loss, language model perplexity, and word error rate.

误差分析贝叶斯误差分类上界语音识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。