arXiv:2509.09067cs.CV2025-09

融合场景与多任务学习,提升物体交互动作识别准确率。

Improvement of Human-Object Interaction Action Recognition Using Scene Information and Multi-Task Learning Approach

  • 引入固定物体位置与交互区域信息增强表征
  • 在真实场景数据集上达99.25%准确率,较基线提升2.75%
  • 适合关注智能监控、人机交互的开发者

近期图卷积神经网络(GCNs)在人体动作识别中表现优异,主要依赖人体骨骼姿态。然而,在检测人-物交互动作时表现不佳,原因在于缺乏有效的场景信息表征及合适的网络结构。为此,本文提出一种结合环境内固定物体信息与多任务学习的方法,以提升人体动作识别性能。为评估该方法,我们从公开环境采集真实数据,构建包含手触固定物体(如ATM取票机、出入站机等)的交互类,以及行走、站立等非交互类的动作数据集。通过引入多任务学习框架与交互区域信息,所提方法在测试中实现了99.25%的识别准确率,相较仅使用人体骨骼姿态的基线模型提升了2.75%。

原文摘要 · Abstract (English)

Recent graph convolutional neural networks (GCNs) have shown high performance in the field of human action recognition by using human skeleton poses. However, it fails to detect human-object interaction cases successfully due to the lack of effective representation of the scene information and appropriate learning architectures. In this context, we propose a methodology to utilize human action recognition performance by considering fixed object information in the environment and following a multi-task learning approach. In order to evaluate the proposed method, we collected real data from public environments and prepared our data set, which includes interaction classes of hands-on fixed objects (e.g., ATM ticketing machines, check-in/out machines, etc.) and non-interaction classes of walking and standing. The multi-task learning approach, along with interaction area information, succeeds in recognizing the studied interaction and non-interaction actions with an accuracy of 99.25%, outperforming the accuracy of the base model using only human skeleton poses by 2.75%.

动作识别多任务学习人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。