arXiv:2608.29611cs.CV2026-08

通过时空表示学习提升无监督动作分割的伪标签质量

See the Change, Keep the Flow: Unsupervised Action Segmentation via Spectral-Temporal Representation Learning

论文配图:See the Change, Keep the Flow: Unsupervised Action Segmentation via Spectral-Temporal Representation Learning
图 1 · 摘自论文原文
  • 设计谱-时序表示框架,用傅里叶基参数化投影器增强变化敏感性
  • 引入时序亲和正则化,保持局部时间结构稳定性,提升伪标签可靠性
  • 在4个数据集上15项指标中13项领先,最大提升7.4点F1

无监督动作分割旨在不依赖动作标注的情况下发现潜在动作类别及其时间组织。基于最优传输(OT)的方法能提供结构化的帧到动作分配,但其伪标签质量受限于传输代价所用的表示空间。本文认为,可靠的OT伪标签需要表示空间同时对动作变化具有判别敏感性,并在局部时间进程中保持一致性。基于此,提出SpecT-OT框架,基于非平衡最优传输伪标签思想,引入谱重参数化投影器(SRP),以固定傅里叶基参数化投影权重并学习可调系数,增强对快速变化特征的建模;同时提出时序亲和正则化(TAR),对成对帧亲和施加距离感知、无需标签的约束,稳定局部时间结构。二者协同生成更具判别性和时间稳定的传输代价,从而提升迭代表示学习中的伪标签可靠性。在四个基准数据集上的实验表明,SpecT-OT在15项指标中有13项达到最优,相较基线在Breakfast数据集上提升4.1点平均重叠率(MoF),在Desktop Assembly上提升7.4点F1。

原文摘要 · Abstract (English)

Unsupervised action segmentation aims to discover latent action categories and their temporal organization without action annotations. Optimal transport-based methods provide structured frame-to-action assignments, however, their pseudo-label quality is fundamentally conditioned on the representation space used to construct the transport cost. We argue that reliable OT pseudo-labeling requires a representation geometry that is simultaneously sensitive to discriminative action changes and coherent along local temporal progressions. Based on this insight, we propose SpecT-OT, a spectral-temporal representation learning framework built upon an unbalanced optimal transport pseudo-labeling concept. SpecT-OT introduces a Spectral Reparameterization Projector (SRP), which parameterizes projector weights with fixed Fourier bases and learnable coefficients to improve the modeling of rapidly varying discriminative features, and Temporal Affinity Regularization (TAR), which imposes distance-aware, label-free constraints on pairwise frame affinities to stabilize local temporal structure. The two components jointly produce more discriminative and temporally stable transport costs, yielding more reliable pseudo-labels for iterative representation learning. Experiments on four benchmarks demonstrate strong performance compared with state-of-the-art methods. SpecT-OT achieves the best results on 13 of 15 metrics, including 4.1-point MoF and 7.4-point F1 gains over the baseline on Breakfast and Desktop Assembly, respectively.

动作分割无监督学习最优传输时空建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。