用视觉语言动作模型实现双臂机器人自动调酒,精准控液配比。
Shake-VLA: Vision-Language-Action Model-Based System for Bimanual Robotic Manipulations and Liquid Mixing
- 基于VLA模型整合视觉、语音与语言生成,驱动双臂机器人完成复杂操作。
- 在嘈杂环境和杂乱场景下,语音与视觉识别成功率分别达93%和91%。
- 支持异常检测与动态调参,适合需高精度液体混合的自动化任务。
本文提出Shake-VLA系统,基于视觉-语言-动作(VLA)模型,实现双臂机器人自动化调制鸡尾酒。系统集成视觉模块用于识别原料瓶及读取标签,语音转文本模块在噪声环境下解析用户指令,语言模型生成特定任务的机器人控制指令。通过力矩传感器精确测量倾倒液体量,确保配料比例准确。系统架构包含检索增强生成(RAG)模块以获取并适配配方,异常检测机制识别原料与配方不一致问题,并由双臂机械手完成灵巧操作。实验表明,各模块表现优异:语音转文本在噪声环境中成功率达93%,视觉模块在杂乱场景中物体与标签检测准确率为91%,异常检测模块成功识别95%的差异,系统整体从配方生成到动作执行的鸡尾酒制备成功率高达100%。
原文摘要 · Abstract (English)
This paper introduces Shake-VLA, a Vision-Language-Action (VLA) model-based system designed to enable bimanual robotic manipulation for automated cocktail preparation. The system integrates a vision module for detecting ingredient bottles and reading labels, a speech-to-text module for interpreting user commands, and a language model to generate task-specific robotic instructions. Force Torque (FT) sensors are employed to precisely measure the quantity of liquid poured, ensuring accuracy in ingredient proportions during the mixing process. The system architecture includes a Retrieval-Augmented Generation (RAG) module for accessing and adapting recipes, an anomaly detection mechanism to address ingredient availability issues, and bimanual robotic arms for dexterous manipulation. Experimental evaluations demonstrated a high success rate across system components, with the speech-to-text module achieving a 93% success rate in noisy environments, the vision module attaining a 91% success rate in object and label detection in cluttered environment, the anomaly module successfully identified 95% of discrepancies between detected ingredients and recipe requirements, and the system achieved an overall success rate of 100% in preparing cocktails, from recipe formulation to action generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。