arXiv:2509.24382cs.CVcs.AI2025-09被引 1

解决指令视频中冗余片段干扰,通过部分对齐提升步骤识别准确率。

REMAP: Regularized Matching and Partial Alignment of Video Embeddings

  • 采用部分最优传输建模语义与时间结构,允许无关帧不匹配
  • 在EgoProceL上提升11.6% F1(+4.45pp),IoU提升19.6%(+4.73pp)
  • 适合处理长时、噪声多的现实场景教学视频理解任务

真实世界中的教学视频通常很长、含有噪声,包含大量背景段落、重复动作和执行变异性,这些并不对应有意义的步骤。我们提出**REMAP**,一种基于*正则化融合部分格罗莫夫-沃瑟斯坦最优传输*的无监督流程学习框架。REMAP放松了平衡传输约束,允许非信息性或冗余帧通过部分传输保持未匹配状态。该方法联合建模语义相似性与时间结构,并引入拉普拉斯平滑与结构正则化,防止退化解对齐并减少背景干扰。我们在大规模第一人称和第三人称基准上评估了REMAP,结果持续优于现有先进方法:在EgoProceL上,F1最高提升11.6%(+4.45pp),IoU提升19.6%(+4.73pp);在ProceL与CrossTask上平均F1提升41%(+17.15pp)。结果表明,部分对齐对处理真实场景中的流程变异性至关重要,REMAP为教学视频理解提供了鲁棒且可扩展的解决方案。

原文摘要 · Abstract (English)

Real-world instructional videos are long, noisy, and often contain extended background segments, repeated actions, and execution variability that do not correspond to meaningful procedural steps. We propose **REMAP**, an unsupervised framework for procedure learning based on *Regularized Fused Partial Gromov-Wasserstein Optimal Transport*. REMAP relaxes balanced transport constraints, allowing non-informative or redundant frames to remain unmatched through partial transport. The formulation jointly models semantic similarity and temporal structure, while incorporating Laplacian-based smoothness and structural regularization to prevent degenerate alignments and reduce background interference. We evaluate REMAP on large-scale egocentric and third-person benchmarks. The method consistently outperforms state-of-the-art approaches, achieving up to **11.6\% (+4.45pp)** F1 and **19.6\% (+4.73pp)** IoU improvements on EgoProceL, and an average **41\% (+17.15pp)** F1 gain on ProceL and CrossTask. These results highlight the importance of partial alignment in handling real-world procedural variability and demonstrate that REMAP provides a robust and scalable approach for instructional video understanding.

视频理解流程识别最优传输

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。