arXiv:2606.13436cs.AI2026-06

提出评估主权概念,揭示弱监督下指标可能只反映标签习惯而非真实能力

Evaluation Sovereignty in Metadata-Driven Classification: A Multi-Track Framework for Weakly Supervised Information Systems

  • 构建多路径评估框架,分离训练与评估标签来源以检验指标可靠性
  • 细粒度分类中微平均F1从0.54骤降至0.03,显示性能严重依赖标签权威
  • 适合关注模型审计、弱监督系统可信度的研究者和工程团队

机器学习评估常被视为中立过程,但在实际信息系统中,评估结果往往受标签生成流程影响。本文不追求提升分类性能,而是考察不同标签授权制度下的评估有效性。在大规模元数据驱动系统中,标签常不完整、不一致或为弱监督。我们提出‘评估主权’——即性能指标对标签权威与监督方式的独立程度,并设计多路径评估框架,系统性地改变训练与评估的标签来源。基于大规模科学元数据的分层多标签分类任务发现,模型在运营(银)评估下表现良好,但在独立(金)评估下性能大幅下降,尤其在细粒度分类中:微平均F1由约0.54降至0.03。值得注意的是,排名类指标仍高于基线,表明模型潜在信号与分类有效性之间存在分歧。这说明常见报告的性能指标可能反映的是与标注流程的一致性,而非真实预测能力。因此,我们重新将评估有效性视为受标签治理塑造的系统级属性,并提供一套适用于弱监督智能系统的可操作审计方法。

原文摘要 · Abstract (English)

Evaluation in machine learning is typically treated as a neutral measurement process. However, in operational information systems, evaluation outcomes are often conditioned by the processes used to generate labels. This paper does not seek to improve classification performance. Instead, it examines the validity of performance measurement under differing label-authority regimes. This issue is particularly relevant in large-scale metadata-driven systems, where labels are often incomplete, inconsistent, or weakly supervised. We introduce evaluation sovereignty, defined as the degree to which performance metrics are independent of label authority and supervision regime, and propose a multi-track evaluation framework that systematically varies training and evaluation label sources. Using hierarchical multi-label classification on large-scale scientific metadata, we demonstrate that models exhibiting strong performance under operational ("silver") evaluation degrade substantially under independent ("gold") evaluation, particularly for fine-grained classification. For example, Micro-F1 decreases from approximately 0.54 to 0.03. Notably, ranking-based metrics remain above baseline, revealing a divergence between latent model signal and classification validity. These findings suggest that commonly reported performance metrics may reflect alignment with labeling processes rather than true predictive capability. We therefore reconceptualize evaluation validity as a system-level property shaped by label governance and provide a practical methodology for auditing intelligent systems operating under weak supervision.

评估主权弱监督模型审计元数据分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。