arXiv:2509.23112cs.RO2025-09被引 5

用力矩传感增强视觉,让机器人更准抓取和翻转瓶子

FTACT: Force Torque aware Action Chunking Transformer for Pick-and-Reorient Bottle Task

  • 结合图像、关节状态与力矩数据,端到端训练抓取策略
  • 在真实机器人上任务成功率显著高于仅用视觉的基线
  • 适合需要精准接触感知的零售场景机器人开发

机械臂在零售环境中的应用日益广泛,但高接触密度的边缘情况仍需人工远程操控。以直立放置的饮料瓶为例,仅靠视觉常无法识别细微接触事件。本文提出一种多模态模仿学习策略,将力矩感知融入动作分块变换器(Action Chunking Transformer),实现图像、关节状态及力/矩信号的联合端到端学习。在Telexistence公司开发的Ghost单臂平台上部署,该方法通过检测并利用按压与放置过程中的接触状态变化,显著提升抓取并翻转瓶子的任务成功率。硬件实验表明,在与仅使用视觉观测的基线对比中表现更优;力矩信号在视觉信息不足的按压与放置阶段尤其有效,验证了交互力作为接触密集型技能补充模态的价值。结果表明,结合现代模仿学习架构与轻量级力矩传感,是推动零售场景机器人规模化落地的可行路径。

原文摘要 · Abstract (English)

Manipulator robots are increasingly being deployed in retail environments, yet contact rich edge cases still trigger costly human teleoperation. A prominent example is upright lying beverage bottles, where purely visual cues are often insufficient to resolve subtle contact events required for precise manipulation. We present a multimodal Imitation Learning policy that augments the Action Chunking Transformer with force and torque sensing, enabling end-to-end learning over images, joint states, and forces and torques. Deployed on Ghost, single-arm platform by Telexistence Inc, our approach improves Pick-and-Reorient bottle task by detecting and exploiting contact transitions during pressing and placement. Hardware experiments demonstrate greater task success compared to baseline matching the observation space of ACT as an ablation and experiments indicate that force and torque signals are beneficial in the press and place phases where visual observability is limited, supporting the use of interaction forces as a complementary modality for contact rich skills. The results suggest a practical path to scaling retail manipulation by combining modern imitation learning architectures with lightweight force and torque sensing.

机器人抓取力觉感知模仿学习多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。