提出新基准Ego4OOD,量化视角视频的分布偏移问题。
Ego4OOD: Rethinking Egocentric Video Domain Generalization via Covariate Shift Scoring
- 用聚类法定义可测量的输入分布差异,分离概念漂移
- 轻量网络在双数据集上超越主流方法,参数更少
- 适合研究视角视频泛化能力的算法开发者
视角视频在分布偏移下的动作识别仍具挑战,源于类内时空变异大、特征分布长尾且动作与环境强相关。现有基准常混淆输入分布变化与概念变化,难以真实评估泛化能力。为此,我们提出基于Ego4D的Ego4OOD基准,强调可量化的输入多样性,通过语义一致的片段级动作类别减少概念漂移。该基准覆盖八个地理域,并引入基于聚类的协变量偏移度量,提供领域难度的定量指标。我们采用一对多二分类训练目标,将多类识别拆解为独立二分类任务,有效缓解特征分布偏移下视觉相似类别的干扰。实验表明,仅两层全连接的轻量模型在Argo1M和Ego4OOD上性能媲美当前最优方法,参数更少且无需额外模态。实证分析揭示协变量偏移度量与识别性能间明确关系,凸显受控基准与量化表征对研究视角视频分布外泛化的关键作用。
原文摘要 · Abstract (English)
Egocentric video action recognition under domain shifts remains challenging due to large intra-class spatio-temporal variability, long-tailed feature distributions, and strong correlations between actions and environments. Existing benchmarks for egocentric domain generalization often conflate covariate shifts with concept shifts, making it difficult to reliably evaluate a model's ability to generalize across input distributions. To address this limitation, we introduce Ego4OOD, a domain generalization benchmark derived from Ego4D that emphasizes measurable covariate diversity while reducing concept shift through semantically coherent, moment-level action categories. Ego4OOD spans eight geographically distinct domains and is accompanied by a clustering-based covariate shift metric that provides a quantitative proxy for domain difficulty. We further leverage a one-vs-all binary training objective that decomposes multi-class action recognition into independent binary classification tasks. This formulation is particularly well-suited for covariate shift by reducing interference between visually similar classes under feature distribution shift. Using this formulation, we show that a lightweight two-layer fully connected network achieves performance competitive with state-of-the-art egocentric domain generalization methods on both Argo1M and Ego4OOD, despite using fewer parameters and no additional modalities. Our empirical analysis demonstrates a clear relationship between measured covariate shift and recognition performance, highlighting the importance of controlled benchmarks and quantitative domain characterization for studying out-of-distribution generalization in egocentric video.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。