arXiv:2411.15557cs.CVcs.AI2024-11

通过语义关系对齐实现跨域动作预测,提升模型泛化能力。

Semantically Guided Action Anticipation

  • 基于语义几何关系构建无监督域适应框架
  • 在四个数据集上平均准确率提升1.94%~5.75%
  • 适合跨域视频理解与动作预测研究者

无监督域适应在跨域知识迁移中仍面临关键挑战。现有方法难以兼顾域不变表征与保留域特定特征,常因绝对坐标对齐导致语义相似但领域差异大的样本被强制靠近。本文提出新思路:不再对齐潜在空间中的绝对位置,而是对齐等价概念间的相对位置关系。通过语言空间中类别标签的语义/几何关系建立无域依赖结构,引导视觉空间中样本组织反映参考类间关系,同时保留域特异性特征。我们在四个多样化的图像与视频数据集上实证该方法优越性,在18个不同适配场景中表现领先:在DomainNet上平均准确率提升+3.32%,GeoPlaces上+5.75%,GeoImnet上+4.77%,EgoExo4D上均类准确率提升+1.94%。

原文摘要 · Abstract (English)

Unsupervised domain adaptation remains a critical challenge in enabling the knowledge transfer of models across unseen domains. Existing methods struggle to balance the need for domain-invariant representations with preserving domain-specific features, which is often due to alignment approaches that impose the projection of samples with similar semantics close in the latent space despite their drastic domain differences. We introduce a novel approach that shifts the focus from aligning representations in absolute coordinates to aligning the relative positioning of equivalent concepts in latent spaces. Our method defines a domain-agnostic structure upon the semantic/geometric relationships between class labels in language space and guides adaptation, ensuring that the organization of samples in visual space reflects reference inter-class relationships while preserving domain-specific characteristics. We empirically demonstrate our method's superiority in domain adaptation tasks across four diverse image and video datasets. Remarkably, we surpass previous works in 18 different adaptation scenarios across four diverse image and video datasets with average accuracy improvements of +3.32% on DomainNet, +5.75% in GeoPlaces, +4.77% on GeoImnet, and +1.94% mean class accuracy improvement on EgoExo4D.

动作预测域适应语义对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。