为制造机器人设计逻辑感知的物体操作知识框架,提升智能助手执行精度。
Towards Logic-Aware Manipulation: A Knowledge Primitive for VLM-Based Assistants in Smart Manufacturing
- 提出八字段物体中心操作逻辑结构τ,显式表达力、轨迹等关键参数。
- 在3D打印机线轴更换任务中,τ条件下的规划质量显著优于传统方法。
- 适用于智能制造中的智能助手系统,支持训练增强与测试推理优化。
现有视觉语言模型在机器人操作中的应用侧重于图像与语言的广泛语义泛化,但通常忽略制造场景中接触密集型操作所需的执行关键参数。本文形式化定义了一种以物体为中心的操作逻辑架构τ,以八字段元组形式封装对象、接触面、轨迹、公差及力/阻抗等信息,作为人机协同、VLM助手与机器人控制器之间的第一类知识信号。我们在协作制造单元中针对3D打印机线轴更换任务实例化τ及小型知识库(KB),并采用近期VLM/LLM规划评估基准中的规划质量指标分析τ条件下的规划表现;同时证明该架构可在训练时支持基于分类标签的数据增强,在测试时实现逻辑感知的检索增强提示,成为智能制造企业中助手系统的核心构建模块。
原文摘要 · Abstract (English)
Existing pipelines for vision-language models (VLMs) in robotic manipulation prioritize broad semantic generalization from images and language, but typically omit execution-critical parameters required for contact-rich actions in manufacturing cells. We formalize an object-centric manipulation-logic schema, serialized as an eight-field tuple τ, which exposes object, interface, trajectory, tolerance, and force/impedance information as a first-class knowledge signal between human operators, VLM-based assistants, and robot controllers. We instantiate τ and a small knowledge base (KB) on a 3D-printer spool-removal task in a collaborative cell, and analyze τ-conditioned VLM planning using plan-quality metrics adapted from recent VLM/LLM planning benchmarks, while demonstrating how the same schema supports taxonomy-tagged data augmentation at training time and logic-aware retrieval-augmented prompting at test time as a building block for assistant systems in smart manufacturing enterprises.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。