arXiv:2603.27197cs.CV2026-03中稿 · CVPR

提出KαLOS算法,解决视觉任务标注一致性评估难题。

K$α$LOS finds Consensus: A Meta-Algorithm for Evaluating Inter-Annotator Agreement in Complex Vision Tasks

  • 先定位后判断,通过空间对应统一标注评估标准
  • 可检测标注者活力、协作集群与定位敏感度等细粒度问题
  • 适用于框、分割、姿态等多种复杂视觉任务

目标检测基准进展停滞,根源不在于模型架构,而在于难以区分模型改进与标注噪声。要重建基准评测可信度,需严格量化标注一致性以保障评估数据可靠性。然而,标准统计指标无法处理视觉任务中的实例对应问题。且新一致性度量缺乏客观真实标签验证,导致依赖不可验证的启发式方法。本文提出KαLOS(KALOS),一种统一元算法,将“先定位”原则推广至标准化数据集质量评估。通过在评估一致性前解决空间对应问题,该框架将复杂的时空分类问题转化为名义可靠性矩阵。不同于以往启发式实现,KαLOS采用数据驱动的严谨配置:通过统计校准定位参数以匹配内在一致性分布,从而泛化到从边界框到体积分割或姿态估计等多种任务。该标准化支持细粒度诊断,包括标注者活力、协作聚类和定位敏感性。为验证方法,我们引入一种新型可控噪声生成器。与以往基于均匀误差假设的验证不同,该测试平台建模了复杂且非各向同性的人员变异。这揭示了度量的性质,并确立了KαLOS作为现代计算机视觉基准中区分信号与噪声的可靠标准。

原文摘要 · Abstract (English)

Progress in object detection benchmarks is stagnating. It is limited not by architectures but by the inability to distinguish model improvements from label noise. To restore trust in benchmarking the field requires rigorous quantification of annotation consistency to ensure the reliability of evaluation data. However, standard statistical metrics fail to handle the instance correspondence problem inherent to vision tasks. Furthermore, validating new agreement metrics remains circular because no objective ground truth for agreement exists. This forces reliance on unverifiable heuristics. We propose K$α$LOS (KALOS), a unified meta-algorithm that generalizes the "Localization First" principle to standardize dataset quality evaluation. By resolving spatial correspondence before assessing agreement, our framework transforms complex spatio-categorical problems into nominal reliability matrices. Unlike prior heuristic implementations, K$α$LOS employs a principled, data-driven configuration; by statistically calibrating the localization parameters to the inherent agreement distribution, it generalizes to diverse tasks ranging from bounding boxes to volumetric segmentation or pose estimation. This standardization enables granular diagnostics beyond a single score. These include annotator vitality, collaboration clustering, and localization sensitivity. To validate this approach, we introduce a novel and empirically derived noise generator. Where prior validations relied on uniform error assumptions, our controllable testbed models complex and non-isotropic human variability. This provides evidence of the metric's properties and establishes K$α$LOS as a robust standard for distinguishing signal from noise in modern computer vision benchmarks.

标注一致性视觉任务数据质量元算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。