arXiv:2608.10718cs.RO2026-08

用自研系统实现全自动叠衣操作,23秒完成一件,精准度高。

TCAM for Autonomous Deformable Manipulation: The RMC2 Champion System for WBCD 2026 Track 4

论文配图:TCAM for Autonomous Deformable Manipulation: The RMC2 Champion System for WBCD 2026 Track 4
图 1 · 摘自论文原文
  • 基于TCAM框架,融合硬件、感知与学习,降低抓取复杂性。
  • 实测每件衣服平均耗时23秒,22/25次达成平整要求。
  • 适合做柔性物体自动化操作的研究者和工程师参考。

本文介绍RMC2团队在WBCD 2026 Track 4:可变形物体操控挑战赛中的冠军解决方案。任务要求机器人从叠放的T恤中取出单件,放置于打印托盘上,对齐领口并平整打印区域,涉及单层分离、柔性搬运、精确定位与接触密集表面调整。比赛强调全自主执行,推动了自主系统的开发。我们基于TCAM(TermiBrain因果动作模型)框架构建了完整自主系统,设计原则是硬件、感知、数据与学习协同降低策略需处理的物理交互复杂度。针对双臂ARX X5平台,定制3D打印夹爪提升单层布料分离可靠性;腕部集成四相机布局,上部广角相机提供任务级上下文,下部RGB相机观测夹爪与布料近距离接触。结合便携式UMI式演示与部署平台上的真实机器人演示,兼顾广泛操控先验与部署特定动力学。TCAM将各组件闭环整合:每条轨迹分析物理因素影响,驱动针对性数据回采与策略微调。策略通过多视角视觉语言模型骨干输出30步末端执行器增量位姿动作块。最终比赛中,系统成功装载25件T恤,平均每次约23秒,其中22件达到所需平整度,夺得Track 4第一名。

原文摘要 · Abstract (English)

This technical report describes the RMC2 Team's champion solution for the WBCD 2026 Track 4: Deformable Manipulation Challenge. The task requires a robot to pick a single T-shirt from a stack, load it onto a printing pallet, align the collar with a target area, and smooth the printing region, a sequence that involves single-layer separation, deformable transport, precise placement, and contact-rich surface adjustment. The competition strongly incentivizes fully autonomous execution, motivating the development of an autonomous solution. We built a fully autonomous system around the TCAM (TermiBrain Causal Action Model) framework, with the design principle that hardware, perception, data, and learning should jointly reduce the physical interaction complexity the policy must handle. A custom 3D-printed gripper designed for single-layer fabric separation improves picking reliability on a dual-arm ARX X5 platform. A wrist-centric four-camera setup pairs upper fisheye cameras for task-level context with lower RGB cameras for close-range gripper-cloth contact observation. We combine portable UMI-style demonstrations with real-robot demonstrations collected on the deployable platform to provide both broad manipulation priors and deployment-specific dynamics. TCAM ties these components into a closed loop: each trajectory is analyzed to identify the physical factors contributing to its outcome, driving targeted data recollection and policy fine-tuning. The policy outputs 30-step end-effector delta-pose action chunks from a multi-view VLA backbone. In the final competition, our system loaded 25 T-shirts at an average of approximately 23 seconds per attempt, with 22 achieving the required surface smoothness, securing first place in Track 4.

机器人操作柔性物体自主系统视觉-动作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。