用代理模型统一多视角多模态知识,提升第一人称视频理解能力
UNIEGO: Proxies as Mediators for Unified Egocentric Video Representation Learning

- 通过代理模型转换不同教师的异构知识到统一空间
- 在三个基准上实现动作识别等任务的最新性能
- 适合需要融合多源信息的第一人称视觉研究者
第一人称视频理解受限于可穿戴摄像头的单一视角、单一路径和单一模型,难以捕捉人类动作的全部丰富性。我们提出一种分层多教师蒸馏框架,训练出UNIEGO——一个统一的第一人称编码器,使用九个教师模型覆盖第一人称-第三人称视角、RGB、深度和骨骼模态,以及四个基础模型。为避免异构教师因架构与特征几何不兼容导致梯度冲突,框架引入代表特定表示的代理模型,将多元知识映射至统一的第一人称空间。第二阶段选择性代理蒸馏(SPD)自适应选取每个样本中正确且置信的代理子集,仅从可靠监督信号蒸馏,抑制错误信息。SPD通过将UNIEGO初始化为代理参数的可学习凸组合,使模型起始位于损失曲面的良好区域。UNIEGO在三个挑战性第一人称-第三人称基准上,于动作识别、视频检索和动作分割任务均达到当前最优表现,显著优于直接多教师蒸馏基线,证明结构化代理中介的知识迁移可生成更丰富、更具判别性的第一人称表示。
原文摘要 · Abstract (English)
Egocentric video understanding is inherently limited by the narrow perspective of wearable cameras: a single viewpoint, a single modality, a single model cannot capture the full richness of human action. We argue that a truly expressive egocentric representation must subsume complementary knowledge across viewpoints, modalities, and foundation model representations, yet remain deployable from egocentric video alone. To this end, we introduce a hierarchical multi-teacher distillation framework that produces UNIEGO, a unified egocentric encoder trained with nine teachers spanning ego-exo viewpoints, RGB, depth, and skeleton modalities, and four foundation models. Rather than distilling directly from heterogeneous teachers whose incompatible architectures and feature geometries induce conflicting gradients, our framework interposes a layer of representation-specific Proxy models that translate diverse teacher knowledge into a homogeneous egocentric space. A second distillation stage, Selective Proxy Distillation (SPD), then adaptively selects, for each training sample, the subset of proxies that are both correct and confident, distilling exclusively from reliable supervision and suppressing erroneous signals. SPD is further stabilized by initializing UNIEGO as a learned convex combination of proxy parameters, placing the unified model in a well-conditioned region of the loss landscape before distillation begins. UNIEGO achieves state-of-the-art performance across three egocentric video understanding tasks - action recognition, video retrieval, and action segmentation on three challenging ego-exo benchmarks, outperforming naive multi-teacher distillation baselines and demonstrating that structured, proxy-mediated knowledge transfer yields richer and more discriminative egocentric representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。