arXiv:2505.18012cs.RO2025-05

用关键点坐标识别复杂装配任务,xLSTM表现优于Transformer。

Classification of assembly tasks combining multiple primitive actions using Transformers and xLSTMs

  • 基于手部关键点坐标,用LSTM、Transformer和xLSTM分类长序列装配任务。
  • Transformer在训练者上准确率达95.0%,xLSTM在新人上达60.8%。
  • xLSTM泛化能力更强,适合跨操作者的协作机器人应用。

人类装配任务的分类在协作机器人中至关重要,有助于保障安全、预判机器人动作并促进其学习。然而,当无法将任务拆分为小的原始动作时,对包含多个原始动作的长序列装配任务进行可靠分类仍具挑战性。本研究提出基于手部关键点坐标对长序列装配任务进行分类,并比较了LSTM、Transformer和新兴的xLSTM模型的表现。实验采用CT基准中提出的HRC场景,涵盖插入、拧紧螺钉、卡扣装配等复合动作。测试数据来自训练者本人及三位新操作者。结果显示,对于训练者,LSTM、Transformer和xLSTM的准确率分别为72.9%、95.0%和93.2%;对于新操作者,准确率分别为43.5%、54.3%和60.8%。LSTM明显落后于另两者。尽管Transformer和xLSTM在训练者上均表现良好,但xLSTM在新操作者上的泛化性能更优。结果表明,xLSTM在该任务中略优于Transformer。

原文摘要 · Abstract (English)

The classification of human-performed assembly tasks is essential in collaborative robotics to ensure safety, anticipate robot actions, and facilitate robot learning. However, achieving reliable classification is challenging when segmenting tasks into smaller primitive actions is unfeasible, requiring us to classify long assembly tasks that encompass multiple primitive actions. In this study, we propose classifying long assembly sequential tasks based on hand landmark coordinates and compare the performance of two well-established classifiers, LSTM and Transformer, as well as a recent model, xLSTM. We used the HRC scenario proposed in the CT benchmark, which includes long assembly tasks that combine actions such as insertions, screw fastenings, and snap fittings. Testing was conducted using sequences gathered from both the human operator who performed the training sequences and three new operators. The testing results of real-padded sequences for the LSTM, Transformer, and xLSTM models was 72.9%, 95.0% and 93.2% for the training operator, and 43.5%, 54.3% and 60.8% for the new operators, respectively. The LSTM model clearly underperformed compared to the other two approaches. As expected, both the Transformer and xLSTM achieved satisfactory results for the operator they were trained on, though the xLSTM model demonstrated better generalization capabilities to new operators. The results clearly show that for this type of classification, the xLSTM model offers a slight edge over Transformers.

装配任务xLSTM关键点泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。