arXiv:2412.09037cs.LG2024-12被引 5

剖析6个主流动作识别数据集的标注模糊问题,揭示模型失败根源。

Beyond Confusion: A Fine-grained Dialectical Examination of Human Activity Recognition Benchmark Datasets

  • 细粒度分析6个数据集,发现多种模型共错区域(IFC)
  • IFC区域中模型准确率不足,暴露标注与采集缺陷
  • 提出三类标注问题分类法,指导未来数据优化

人类活动识别(HAR)领域的机器学习研究依赖公开数据集取得显著进展。然而,多数研究仅关注统计指标,忽视负样本细节。尽管近期模型如Transformer已在基准测试中表现有限,但其同类方法在类似任务上可实现接近100%准确率,引发对当前方法局限性的质疑。本文针对六款主流HAR基准数据集开展细粒度检查,发现部分数据中所有六种先进机器学习方法均无法正确分类,该现象称为交叉错误分类(IFC)。对IFC区域的分析揭示多个深层问题:标注模糊、采集过程异常及状态转换期对齐偏差。本文贡献包括量化并描述标注模糊性,提出用于数据修复的三元分类掩码,并强调未来数据收集应改进的方向。

原文摘要 · Abstract (English)

The research of machine learning (ML) algorithms for human activity recognition (HAR) has made significant progress with publicly available datasets. However, most research prioritizes statistical metrics over examining negative sample details. While recent models like transformers have been applied to HAR datasets with limited success from the benchmark metrics, their counterparts have effectively solved problems on similar levels with near 100% accuracy. This raises questions about the limitations of current approaches. This paper aims to address these open questions by conducting a fine-grained inspection of six popular HAR benchmark datasets. We identified for some parts of the data, none of the six chosen state-of-the-art ML methods can correctly classify, denoted as the intersect of false classifications (IFC). Analysis of the IFC reveals several underlying problems, including ambiguous annotations, irregularities during recording execution, and misaligned transition periods. We contribute to the field by quantifying and characterizing annotated data ambiguities, providing a trinary categorization mask for dataset patching, and stressing potential improvements for future data collections.

动作识别数据质量标注模糊模型诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。