用特征引导的结构化最优传输,解决无监督动作分割中的局部信息丢失和伪标签偏差问题。
FIS-OT: Feature-Induced Optimal Transport for Unsupervised Action Segmentation

- 通过特征增强生成器捕捉局部一致性,避免依赖噪声伪标签
- 结合固定时序主干与动态特征相似性,确保动作边界连续性
- 循环优化机制使局部特征学习与全局结构对齐,适合复杂动作分割任务
无监督动作分割是一项挑战性任务,需在无标签视频中识别动作类别与边界。现有最优传输方法依赖全局约束,忽视局部信息;且其架构易受伪标签偏差影响,在训练初期便学习噪声。为此,本文提出FIS-OT——一种特征诱导的结构化最优传输框架。首先,引入特征增强生成器(FEG)模块,利用三元组损失捕捉局部一致性,提供独立于伪标签的鲁棒监督。其次,提出特征诱导残差结构先验,融合固定时序主干与动态特征相似性,确保时间连续性并适应复杂动作结构。最后,建立循环优化流程,实现局部特征学习与全局结构对齐的协同优化。在三个数据集上的大量实验验证了方法的有效性。
原文摘要 · Abstract (English)
Unsupervised action segmentation is a challenging task. It involves finding action categories and boundaries in videos without labels. Existing Optimal Transport (OT) methods use global constraints. This causes them to overlook the use of local information. Furthermore, existing Optimal transport architectures are prone to confirmation bias because they overly trust the pseudo-labels they generate. This causes models to learn from noise in the early training stages. To address these issues, we propose FIS-OT. It is a novel Feature-Induced Structured Optimal Transport framework. First, we introduce a Feature Enhanced Generator (FEG) module. It serves as an internal regularizer. By using triplet loss, FEG captures local consistency. It provides robust supervision that is independent of noisy pseudo-labels. Second, we propose a Feature-Induced Residual Structural Prior. This combines a fixed temporal backbone with dynamic feature similarities. This design ensures temporal continuity. It also allows the solver to adapt to complex action structures. Finally, we establish a cyclic optimization loop. This aligns local feature learning with global structural alignment. Extensive experiments on the three datasets show the effectiveness of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。