用大模型理解用户自然语言指令,实现影视拍摄机器人动作自动编程
Understanding Generative AI in Robot Logic Parametrization
- 通过大模型将用户口语化拍摄需求转化为机器人可执行的动作参数
- 实验证明能准确捕捉导演意图并生成对应机械臂运动指令
- 适合影视制作、交互式机器人开发等需要快速编程的场景
利用生成式AI(如大语言模型)在机器人中进行语言理解,为语言驱动的机器人用户自定义开发(EUD)带来新可能。尽管设计空间广阔,但如何构建机器人程序逻辑仍不明确。本文以电影摄制为例,探讨摄影师通过自然语言表达拍摄意图,由大模型捕获并转化为机器人机械臂的低层运动参数。研究分析了大模型在迭代程序开发中理解用户意图,并将自然语言映射到预定义跨模态数据的能力。最后提出未来可拓展至更多领域,支持语言驱动的机器人摄像导航。
原文摘要 · Abstract (English)
Leveraging generative AI (for example, Large Language Models) for language understanding within robotics opens up possibilities for LLM-driven robot end-user development (EUD). Despite the numerous design opportunities it provides, little is understood about how this technology can be utilized when constructing robot program logic. In this paper, we outline the background in capturing natural language end-user intent and summarize previous use cases of LLMs within EUD. Taking the context of filmmaking as an example, we explore how a cinematography practitioner's intent to film a certain scene can be articulated using natural language, captured by an LLM, and further parametrized as low-level robot arm movement. We explore the capabilities of an LLM interpreting end-user intent and mapping natural language to predefined, cross-modal data in the process of iterative program development. We conclude by suggesting future opportunities for domain exploration beyond cinematography to support language-driven robotic camera navigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。