arXiv:2507.07745cs.RO2025-07

用大模型分析采摘动作,自动识别和分割基础动作

On the capabilities of LLMs for classifying and segmenting time series of fruit picking motions into primitive actions

  • 用大模型替代传统方法,直接处理运动数据分段
  • 在机械臂采摘数据上验证,能准确识别预设动作类别
  • 适合希望快速部署动作理解系统的机器人研发者

尽管大语言模型(LLMs)进入人类社会时间不长,但已显著改变我们应对日常认知挑战的方式。从优化语言交流到辅助重要决策,像ChatGPT这样的模型正逐步承担越来越多的思维任务,减轻认知负担。在示范学习(LbD)中,将复杂动作分解为基本动作(如推、拉、扭转等)被视为任务编码的关键步骤。本文研究了大模型在该任务中的能力,针对水果采摘操作中预定义的一组基础动作进行分类与分割。通过使用大模型而非简单的监督学习或解析方法,旨在提升方法在真实场景中的可应用性与可部署性。研究比较了三种不同的微调策略,在基于UR10e机械臂采集的触觉运动数据集上进行评估。

原文摘要 · Abstract (English)

Despite their recent introduction to human society, Large Language Models (LLMs) have significantly affected the way we tackle mental challenges in our everyday lives. From optimizing our linguistic communication to assisting us in making important decisions, LLMs, such as ChatGPT, are notably reducing our cognitive load by gradually taking on an increasing share of our mental activities. In the context of Learning by Demonstration (LbD), classifying and segmenting complex motions into primitive actions, such as pushing, pulling, twisting etc, is considered to be a key-step towards encoding a task. In this work, we investigate the capabilities of LLMs to undertake this task, considering a finite set of predefined primitive actions found in fruit picking operations. By utilizing LLMs instead of simple supervised learning or analytic methods, we aim at making the method easily applicable and deployable in a real-life scenario. Three different fine-tuning approaches are investigated, compared on datasets captured kinesthetically, using a UR10e robot, during a fruit-picking scenario.

动作识别大模型机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。