arXiv:2509.11109cs.RO2025-09

用小波变换增强视觉与物理动态对齐,提升双臂机器人操作成功率。

FEWT: Frequency-Enhanced Wavelet-based Transformer for Multimodal Wheeled Bimanual Manipulation

  • 融合小波变换与多尺度注意力,显式对齐时空特征。
  • 在模拟插入任务中成功率显著高于传统Transformer基线。
  • 集成自研触觉织物,适配真实场景微动态扰动。

具身智能将物理世界与信息空间连接,机器人通过模仿学习展现出巨大潜力。本研究采用自建轮式双臂机器人平台,结合外骨骼式遥操作系统,实现直观远程操控并高效采集类人动作数据。为解决空间视觉语义与局部高频物理动态之间的表征不匹配问题,提出轻量级频率对齐的模仿学习框架——频增强小波变换变压器(FEWT)。FEWT包含两个核心模块:频增强高效多尺度注意力(FE-EMA)与时间序列离散小波变换(TS-DWT),显式提取并对齐多尺度特征,提升空间视觉表示与时频校准间的兼容性。关键地,为支持真实部署,该框架进一步扩展为多模态系统,通过无缝集成自研智能触觉织物(STF)传感器于末端执行器,提供局部接触应力信息,补充本体感觉与底盘运动线索,构建共享多模态表征。实验表明,核心FEWT架构在模拟双臂插入任务最困难阶段显著提升成功率,优于广泛使用的动作分块变压器基线。此外,在复杂真实世界的移动与桌面操作任务中,完整STF增强系统有效适应微观动态扰动,带来显著性能提升。

原文摘要 · Abstract (English)

Embodied intelligence bridges the physical world and information spaces, with robots demonstrating immense potential through imitation learning algorithms. In this study, a custom-built wheeled bimanual robotic platform equipped with an exoskeleton-style teleoperation system was utilized to realize intuitive remote manipulation and the efficient collection of anthropomorphic action data. To overcome the representation mismatch between spatial visual semantics and localized high-frequency physical dynamics, we propose a lightweight frequency-aligned imitation-learning framework, termed the Frequency-Enhanced Wavelet-based Transformer (FEWT). FEWT integrates two primary modules: Frequency-Enhanced Efficient Multi-Scale Attention (FE-EMA) and Time-Series Discrete Wavelet Transform (TS-DWT) to explicitly extract and align multi-scale features, improving the compatibility between spatial visual representations and temporal-frequency recalibration. Crucially, for real-world deployment, this framework is further extended into a multimodal system by seamlessly integrating a self-developed Smart Tactile Fabric (STF) sensor into the physical end-effectors, providing local contact-stress information that complements proprioceptive and chassis-motion cues in the shared multimodal representation. Experimental evaluations demonstrate that the core FEWT architecture significantly improves the success rate over the widely used Action Chunking with Transformers baseline, particularly during the most challenging phases of simulated bimanual insertion tasks. Furthermore, in complex real-world mobile and desktop manipulation tasks, the full STF-enhanced system effectively adapts to microscopic dynamic perturbations, yielding substantial performance enhancements.

机器人操作多模态感知小波变换模仿学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。