arXiv:2509.21633cs.AI2025-09被引 1

新评估指标让模糊标签分类更公平,支持语义相似但不完全匹配的预测

Semantic F1 Scores: Fair Evaluation Under Fuzzy Class Boundaries

  • 用标签相似矩阵计算软精度和软召回,避免传统方法对相近标签一概判错
  • 在真实与合成数据上验证,比传统F1更贴近人类判断和实际应用效果
  • 无需严格分类体系,适用于有边界模糊或标注分歧的各类任务

我们提出语义F1分数,一种用于主观或模糊多标签分类的新评估指标,通过量化预测标签与真实标签之间的语义相关性来衡量性能。不同于传统F1将语义相关但不完全相同的标签视为完全错误,语义F1引入标签相似矩阵,计算类比于软精度和软召回的分数,并由此得出最终得分。与现有基于相似度的指标不同,我们的两步精度-召回公式可在不丢弃标签或强制不相似标签匹配的前提下,比较任意大小的标签集合。通过给予语义相关但不完全一致标签部分得分,语义F1更真实反映存在人类分歧或类别边界模糊的领域特征。它承认类别重叠、标注不一致,并指出基于相似预测的下游决策会产生相似结果。通过理论论证和在合成及真实数据上的广泛实证验证,我们证明语义F1具有更高的可解释性和生态效度。由于仅需一个领域适配的相似矩阵,且对误设具有鲁棒性,无需依赖刚性本体,因此可跨任务与模态通用。

原文摘要 · Abstract (English)

We propose Semantic F1 Scores, novel evaluation metrics for subjective or fuzzy multi-label classification that quantify semantic relatedness between predicted and gold labels. Unlike the conventional F1 metrics that treat semantically related predictions as complete failures, Semantic F1 incorporates a label similarity matrix to compute soft precision-like and recall-like scores, from which the Semantic F1 scores are derived. Unlike existing similarity-based metrics, our novel two-step precision-recall formulation enables the comparison of label sets of arbitrary sizes without discarding labels or forcing matches between dissimilar labels. By granting partial credit for semantically related but nonidentical labels, Semantic F1 better reflects the realities of domains marked by human disagreement or fuzzy category boundaries. In this way, it provides fairer evaluations: it recognizes that categories overlap, that annotators disagree, and that downstream decisions based on similar predictions lead to similar outcomes. Through theoretical justification and extensive empirical validation on synthetic and real data, we show that Semantic F1 demonstrates greater interpretability and ecological validity. Because it requires only a domain-appropriate similarity matrix, which is robust to misspecification, and not a rigid ontology, it is applicable across tasks and modalities.

评估指标多标签分类语义相似

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。