人类比AI更擅长在复杂条件下识别动作,因依赖关键视觉线索。
Human-AI Divergence in Ego-centric Action Recognition under Spatial and Spatiotemporal Manipulations
- 用最小可识别区域分析人与AI的识别差异
- 人类对缺失关键线索敏感,模型则渐进退化且误判更自信
- 适合关注人机认知差异与鲁棒性设计的研究者
人类在低分辨率、遮挡和视觉杂乱等真实场景中始终优于当前先进AI模型。本文通过大规模对比实验研究了第一人称动作识别中的“人-机差距”,使用最小可识别识别区域(MIRCs)作为评估基准,即完成可靠识别所需的最小空间或时空范围。基于36段EPIC KITCHENS视频构建的Epic ReduAct数据集,包含多级空间缩减与时间打乱条件。实验招募超3000名参与者,使用Side4Video模型进行性能对比。分析结合定量指标(平均缩减率、识别差距)与定性分析,涵盖高/中/低层视觉特征及动作的时间特性,将动作分为低时间性动作(LTA)与高时间性动作(HTA)。结果表明:人类在从MIRC到subMIRC转换时表现骤降,依赖稀疏但语义关键的线索(如手物交互);而模型退化更平缓,常依赖上下文与中低层特征,甚至在空间缩减下误判信心上升。时间上,人类在保留关键空间线索时仍能抗时间打乱,模型却普遍对时间破坏不敏感,显示类别依赖的时间敏感性差异。
原文摘要 · Abstract (English)
Humans consistently outperform state-of-the-art AI models in action recognition, particularly in challenging real-world conditions involving low resolution, occlusion, and visual clutter. Understanding the sources of this performance gap is essential for developing more robust and human-aligned models. In this paper, we present a large-scale human-AI comparative study of egocentric action recognition using Minimal Identifiable Recognition Crops (MIRCs), defined as the smallest spatial or spatiotemporal regions sufficient for reliable human recognition. We used our previously introduced, Epic ReduAct, a systematically spatially reduced and temporally scrambled dataset derived from 36 EPIC KITCHENS videos, spanning multiple spatial reduction levels and temporal conditions. Recognition performance is evaluated using over 3,000 human participants and the Side4Video model. Our analysis combines quantitative metrics, Average Reduction Rate and Recognition Gap, with qualitative analyses of spatial (high-, mid-, and low-level visual features) and spatiotemporal factors, including a categorisation of actions into Low Temporal Actions (LTA) and High Temporal Actions (HTA). Results show that human performance exhibits sharp declines when transitioning from MIRCs to subMIRCs, reflecting a strong reliance on sparse, semantically critical cues such as hand-object interactions. In contrast, the model degrades more gradually and often relies on contextual and mid- to low-level features, sometimes even exhibiting increased confidence under spatial reduction. Temporally, humans remain robust to scrambling when key spatial cues are preserved, whereas the model often shows insensitivity to temporal disruption, revealing class-dependent temporal sensitivities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。