用自然语言控制机器人抓握力度,实现更智能的机械操作。
Bi-LAT: Bilateral Control-Based Imitation Learning via Natural Language and Action Chunking with Transformers

- 通过双端控制融合语言与动作片段,动态调节力矩
- 在杯叠和拧海绵任务中精准复现指令力值
- 适合需要精细力控的人机协作场景
我们提出Bi-LAT,一种将双边控制与自然语言处理结合的模仿学习框架,用于实现机器人操作中的精确力调节。该框架在主从遥操作中利用关节位置、速度和力矩数据,同时整合视觉与语言线索,动态调整施加的力。通过基于多模态Transformer的模型编码“轻轻抓取杯子”或“用力拧海绵”等人类指令,Bi-LAT能够学习区分真实任务中细微的力控需求。我们在单手杯叠场景和双手拧海绵任务中验证了其性能,实验表明,当使用SigLIP作为语言编码器时,Bi-LAT能有效复现指定力水平,显著提升力控精度。研究结果证明,在模仿学习中融入自然语言提示具有潜力,为更直观、自适应的人机交互铺平道路。
原文摘要 · Abstract (English)
We present Bi-LAT, a novel imitation learning framework that unifies bilateral control with natural language processing to achieve precise force modulation in robotic manipulation. Bi-LAT leverages joint position, velocity, and torque data from leader-follower teleoperation while also integrating visual and linguistic cues to dynamically adjust applied force. By encoding human instructions such as "softly grasp the cup" or "strongly twist the sponge" through a multimodal Transformer-based model, Bi-LAT learns to distinguish nuanced force requirements in real-world tasks. We demonstrate Bi-LAT's performance in (1) unimanual cup-stacking scenario where the robot accurately modulates grasp force based on language commands, and (2) bimanual sponge-twisting task that requires coordinated force control. Experimental results show that Bi-LAT effectively reproduces the instructed force levels, particularly when incorporating SigLIP among tested language encoders. Our findings demonstrate the potential of integrating natural language cues into imitation learning, paving the way for more intuitive and adaptive human-robot interaction. For additional material, please visit: https://mertcookimg.github.io/bi-lat/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。