arXiv:2601.16211cs.CVcs.AI2026-01

解决零样本动作识别中模型依赖物体误判的问题

Why Can't I Open My Drawer? Mitigating Object-Driven Shortcuts in Zero-Shot Compositional Action Recognition

  • 引入双组件正则化,抑制物体类别导致的捷径学习
  • 在Sth-com和EK100-com上显著降低捷径诊断指标
  • 适合研究零样本泛化与动作理解的开发者

零样本组合动作识别(ZS-CAR)需识别由已见语义单元构成的新动词-物体组合。本文针对关键缺陷:模型通过物体类别捷径(即依赖标签物体类)而非时间证据预测动词。分析表明,稀疏组合监督与动词-物体学习不对称会助长此类捷径。诊断指标显示,现有方法过度拟合训练共现模式,忽视时间动词线索,导致对未见组合泛化能力弱。为此提出鲁棒组合表征框架RCORE,包含两项设计:共现先验正则化(CPR)通过将高频共现视为硬负例,显式监督未见组合;时序顺序正则化(TORC)强制模型学习时序敏感的动词表征。在Sth-com与EK100-com数据集上,RCORE有效降低捷径诊断值,提升组合泛化性能。

原文摘要 · Abstract (English)

Zero-Shot Compositional Action Recognition (ZS-CAR) requires recognizing novel verb-object combinations composed of previously observed primitives. In this work, we tackle a key failure mode: models predict verbs via object-driven shortcuts (i.e., relying on the labeled object class) rather than temporal evidence. We argue that sparse compositional supervision and verb-object learning asymmetry can promote object-driven shortcut learning. Our analysis with proposed diagnostic metrics shows that existing methods overfit to training co-occurrence patterns and underuse temporal verb cues, resulting in weak generalization to unseen compositions. To address object-driven shortcuts, we propose Robust COmpositional REpresentations (RCORE) with two components. Co-occurrence Prior Regularization (CPR) adds explicit supervision for unseen compositions and regularizes the model against frequent co-occurrence priors by treating them as hard negatives. Temporal Order Regularization for Composition (TORC) enforces temporal-order sensitivity to learn temporally grounded verb representations. Across Sth-com and EK100-com, RCORE reduces shortcut diagnostics and consequently improves compositional generalization.

动作识别零样本捷径学习时序建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。