arXiv:2504.02512cs.CVcs.AI2025-04IJCV

让动作分割模型适应未见过的拍摄视角,提升跨视角泛化能力。

Towards Generalizing Temporal Action Segmentation to Unseen Views

  • 在序列和片段级别构建共享表征,降低视角差异影响。
  • 在未见视角上实现12.8%的F1@50提升,对自上而下视角提升54%。
  • 适用于跨视角动作识别,尤其适合真实场景中多视角部署。

尽管时间动作分割已取得显著进展,但模型在未见视角下的泛化能力仍待解决。为此,本文定义了未见视角动作分割评估协议,即测试时的摄像机视角在训练阶段不可见,包括从俯视/正面视角转为侧视,甚至从外源视角(exocentric)转为第一人称视角(egocentric)。针对该挑战,提出一种新方法:通过引入序列级损失与动作级损失,在视频和动作两个层级建立一致的共享表示,以缓解视角差异带来的影响。在Assembly101、IkeaASM和EgoExoLearn数据集上的实验表明,该方法在未见外源视角上达到F1@50提升12.8%,在未见第一人称视角上更是实现54%的显著改进。

原文摘要 · Abstract (English)

While there has been substantial progress in temporal action segmentation, the challenge to generalize to unseen views remains unaddressed. Hence, we define a protocol for unseen view action segmentation where camera views for evaluating the model are unavailable during training. This includes changing from top-frontal views to a side view or even more challenging from exocentric to egocentric views. Furthermore, we present an approach for temporal action segmentation that tackles this challenge. Our approach leverages a shared representation at both the sequence and segment levels to reduce the impact of view differences during training. We achieve this by introducing a sequence loss and an action loss, which together facilitate consistent video and action representations across different views. The evaluation on the Assembly101, IkeaASM, and EgoExoLearn datasets demonstrate significant improvements, with a 12.8% increase in F1@50 for unseen exocentric views and a substantial 54% improvement for unseen egocentric views.

动作分割跨视角泛化能力第一人称

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。