arXiv:2608.24043cs.CV2026-08

无需标注,自监督分割长时序施工视频中的细粒度动作阶段。

ConsensusTAS: Self-Supervised Temporal Action Segmentation for Long-Horizon Construction Videos

论文配图:ConsensusTAS: Self-Supervised Temporal Action Segmentation for Long-Horizon Construction Videos
图 1 · 摘自论文原文
  • 利用候选分割结果的内在一致性自动生成标签,实现无监督动作分段。
  • 在GTEA、Breakfast等数据集上F1@10最高达73.08,真实施工视频中成功识别砌砖全过程。
  • 仅需CPU运行,适合移动机器人平台部署,实用性强。

识别连续施工活动对人机协同至关重要;例如,机器人可理解工人当前及下一步动作,及时提供工具递送或物理支持。然而,现有研究多局限于分类活动类别(如攀爬、举重、行走),未能识别长时序序列中的细粒度动作转换。这一问题难解,因在长施工视频中标注动作时间边界耗时费力。本研究提出ConsensusTAS,一种无需标签的自监督学习方法,通过挖掘候选分割结果的内部一致性,将连续视频流划分为不同动作阶段。我们在三个公开数据集上评估该算法,表现优于现有最先进方法:在GTEA上取得F1@10为73.08,在Breakfast上为64.33,在Assembly101静态摄像头视频上为F1@50为33.50。此外,在真实施工视频测试中,后验评估显示模型成功识别并分割了砌砖复合动作的多个子阶段,包括抹灰、放砖、按压和对齐。相较于依赖计算密集型大视觉语言模型的其他时序动作分割模型,本方法可在CPU上运行,为视频监控与移动机器人平台上的人机协作提供实用价值。

原文摘要 · Abstract (English)

Recognizing sequential construction activities is important for collaborative human-robot work; for example, robots are able to understand workers' current and upcoming actions and provide timely tool delivery or physical support. However, despite extensive research on construction worker activity recognition, existing studies have been limited to classifying activity categories, such as climbing, lifting, and walking, instead of recognizing fine-grained activity transitions from long-horizon sequences. Addressing this problem is challenging because annotating action temporal boundaries in long construction videos is time-consuming. In this study, we propose ConsensusTAS, a label-free, self-supervised learning approach to segment continuous video streams into distinct activity phases by exploiting the internal consensus of candidate segmentations. We evaluated our algorithm on three public datasets, where it outperformed state-of-the-art methods, achieving an F1@10 of 73.08 on GTEA, an F1@10 of 64.33 on Breakfast, and an F1@50 of 33.50 on static-camera videos from Assembly101. We also tested it on real-world construction videos, where post-hoc evaluation showed that the model successfully recognized and segmented actions within the composite activity of bricklaying, such as spreading mortar on a brick, placing the brick, pressing, and aligning. Compared with other temporal action segmentation models that require computationally intensive large vision-language models, our method can run on a CPU, which provides practical value for video surveillance and human-robot collaboration on mobile robotic platforms.

动作分割自监督施工视频人机协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。