用双机械臂磁控微机器人,实现高精度远程操作。
Mag-VLA: Vision-Language-Action Model for Bimanual Magnetically Actuated Microrobot Manipulation

- 基于视觉语言动作模型,通过双臂协同生成连贯控制指令。
- 真实实验中任务成功率最高达90%,随难度增加降至50%。
- 适合微创医疗、微纳操作等需要精准远程操控的场景。
磁驱动微机器人作为无接触、无线操控工具,在微尺度上具有微创应用潜力。但其控制因间接驱动、感知受限及非线性磁相互作用而困难重重。本文提出Mag-VLA,一种用于双机械臂磁控微机器人灵巧操作的视觉-语言-动作(VLA)模型。双臂协作可实现单臂难以完成的微机器人重定向,但也带来耦合控制挑战。模型采用Qwen2.5-VL-7B骨干网络,结合低秩适配(LoRA)处理视觉输入与语言指令,预测动作。为捕捉任务进展,引入运动感知阶段分类器与阶段条件动作分块变换器(ACT)解码器,实现时序一致的多步控制。构建了包含三种任务配置的遥操作数据集。消融实验证明ACT解码器显著优于其他生成式动作头。真实机器人实验显示,所有任务的接近成功率均达90%,运输成功率分别为80%、70%和50%,随任务难度上升而下降。结果表明,层次化VLA建模为磁控微机器人操作提供了有效框架。
原文摘要 · Abstract (English)
Magnetically actuated microrobots have been used as wireless, non-contact manipulation tools at microscales, making them promising for minimally invasive applications. However, their control remains challenging due to indirect actuation, limited sensing, and nonlinear magnetic interactions. In this work, we propose Mag-VLA, a vision-language-action (VLA) model for dexterous magnetic microrobot manipulation using two robotic arms with mounted magnets for dynamic magnetic-field construction. Bimanual coordination enables capabilities such as microrobot reorientation that are difficult or infeasible with a single arm, but it also introduces coupled control challenges, as the policy must generate coordinated trajectories for both actuators within a shared workspace. Our framework adapts a Qwen2.5-VL-7B backbone using Low-Rank Adaptation (LoRA) to process visual observations and language instructions for action prediction. To capture task progression, we introduce a motion-aware phase classifier and a phase-conditioned Action Chunking Transformer (ACT) decoder for temporally coherent multi-step control. We further construct a teleoperated magnetic microrobot manipulation dataset covering three task configurations. Ablation studies show that the ACT-based decoder substantially outperforms alternative generative action heads. In real-robot experiments, Mag-VLA achieves a 90% approach success rate across all tasks and transport success rates of 80%, 70%, and 50% as task difficulty increases. These results demonstrate that hierarchical VLA modeling provides a promising framework for magnetic microrobot manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。