用小模型实现实验室机器人高效模仿学习,节省算力还靠谱。
Compact Task-Aligned Imitation Learning for Laboratory Automation
- 用轻量适配器对齐视觉与语言模型,结合扩散模型生成动作。
- 全模型参数少于500M,低显存显卡也能运行,三任务平均成功率86.6%。
- 适合资源受限的实验室自动化场景,尤其适合想快速部署的工程团队。
机器人实验室自动化传统上依赖精心设计的运动流水线和特定硬件接口,导致设计成本高且灵活性差。尽管近期模仿学习技术可生成通用机器人行为,但其庞大的模型规模通常需要高性能计算资源,限制了在实际实验环境中的应用。本研究提出一种面向实验室自动化的紧凑型模仿学习框架,采用小型基础模型。所提方法TVF-DiT通过紧凑适配器将自监督视觉基础模型与视觉语言模型对齐,并与基于扩散变换器的动作专家融合。整个模型参数少于500M,可在低显存GPU上推理。在三个真实实验任务——试管清洗、试管排列和粉末转移——上的实验表明,平均成功率达86.6%,显著优于其他轻量化基线。此外,详细的任务提示提升了视觉-语言对齐效果和任务性能。结果表明,当通过适当对齐与扩散策略整合时,小型基础模型可在计算资源有限的情况下有效支持实用的实验室自动化。
原文摘要 · Abstract (English)
Robotic laboratory automation has traditionally relied on carefully engineered motion pipelines and task-specific hardware interfaces, resulting in high design cost and limited flexibility. While recent imitation learning techniques can generate general robot behaviors, their large model sizes often require high-performance computational resources, limiting applicability in practical laboratory environments. In this study, we propose a compact imitation learning framework for laboratory automation using small foundation models. The proposed method, TVF-DiT, aligns a self-supervised vision foundation model with a vision-language model through a compact adapter, and integrates them with a Diffusion Transformer-based action expert. The entire model consists of fewer than 500M parameters, enabling inference on low-VRAM GPUs. Experiments on three real-world laboratory tasks - test tube cleaning, test tube arrangement, and powder transfer - demonstrate an average success rate of 86.6%, significantly outperforming alternative lightweight baselines. Furthermore, detailed task prompts improve vision-language alignment and task performance. These results indicate that small foundation models, when properly aligned and integrated with diffusion-based policy learning, can effectively support practical laboratory automation with limited computational resources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。