arXiv:2507.00756cs.CVcs.RO2025-07中稿 · IROS25, Hangzhou, …被引 3

提出新框架,让模型能自动识别和分割从未见过的人体动作。

Towards Open-World Human Action Segmentation Using Graph Convolutional Networks

  • 用图卷积网络建模动作时空特征,提升泛化能力。
  • 在两个数据集上实现16.9%和34.6%的性能提升。
  • 适合做开放世界动作识别的科研与工程人员参考。

人体-物体交互分割是日常活动理解的核心任务,在辅助机器人、医疗健康和自动驾驶中具有重要意义。现有基于学习的方法在封闭世界场景下表现良好,但难以适应开放世界中不断涌现的新动作。由于人类活动动态多样,难以穷举所有动作类别进行训练,亟需无需人工标注即可检测和分割分布外动作的模型。为此,本文首次正式定义开放世界动作分割问题,并提出一个结构化框架:引入增强型金字塔图卷积网络(EPGCN)与新型解码器,实现鲁棒的时空特征上采样;采用Mixup训练合成分布外数据,避免依赖人工标注;设计新型时间聚类损失,使同类动作聚集、异类样本分离。在双人动作和双手与物体(H2O)两个挑战性数据集上评估,实验结果表明,该框架在多个开放集评价指标上显著优于当前最优模型,开放集分割(F1@50)相对提升16.9%,分布外检测性能(AUROC)提升34.6%。通过深入消融实验,确定了最优框架配置。

原文摘要 · Abstract (English)

Human-object interaction segmentation is a fundamental task of daily activity understanding, which plays a crucial role in applications such as assistive robotics, healthcare, and autonomous systems. Most existing learning-based methods excel in closed-world action segmentation, they struggle to generalize to open-world scenarios where novel actions emerge. Collecting exhaustive action categories for training is impractical due to the dynamic diversity of human activities, necessitating models that detect and segment out-of-distribution actions without manual annotation. To address this issue, we formally define the open-world action segmentation problem and propose a structured framework for detecting and segmenting unseen actions. Our framework introduces three key innovations: 1) an Enhanced Pyramid Graph Convolutional Network (EPGCN) with a novel decoder module for robust spatiotemporal feature upsampling. 2) Mixup-based training to synthesize out-of-distribution data, eliminating reliance on manual annotations. 3) A novel Temporal Clustering loss that groups in-distribution actions while distancing out-of-distribution samples. We evaluate our framework on two challenging human-object interaction recognition datasets: Bimanual Actions and 2 Hands and Object (H2O) datasets. Experimental results demonstrate significant improvements over state-of-the-art action segmentation models across multiple open-set evaluation metrics, achieving 16.9% and 34.6% relative gains in open-set segmentation (F1@50) and out-of-distribution detection performances (AUROC), respectively. Additionally, we conduct an in-depth ablation study to assess the impact of each proposed component, identifying the optimal framework configuration for open-world action segmentation.

动作分割开放世界图神经网络零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。