arXiv:2602.19764cs.RO2026-02

融合视觉力觉等多模态信号,提升机器人灵巧操作能力

Towards Dexterous Embodied Manipulation via Deep Multi-Sensory Fusion and Sparse Expert Scaling

  • 用扩散Transformer统一处理图像、深度和六维力信号
  • 仿真与真实场景成功率分别达83.2%和72.5%,性能领先
  • 稀疏专家架构兼顾模型容量与实时性,适合复杂物理任务

实现灵巧的具身操作需要深度融合异构多模态感知输入。然而,当前以视觉为中心的方法常忽略复杂任务中至关重要的力反馈与几何信息。本文提出DeMUSE框架,采用扩散Transformer将RGB、深度和六轴力信号整合为统一序列流。通过自适应模态归一化(AdaMN)重新校准模态感知特征,缓解表征不平衡,调和多感官信号的异构分布。为实现高效扩展,采用共享专家的稀疏专家混合(MoE)结构,在增加物理先验容量的同时保持实时控制所需的低推理延迟。联合去噪目标同步生成环境演化与动作序列,确保物理一致性。在仿真与真实世界测试中分别取得83.2%和72.5%的成功率,验证了深层多模态融合对复杂物理交互的必要性。

原文摘要 · Abstract (English)

Realizing dexterous embodied manipulation necessitates the deep integration of heterogeneous multimodal sensory inputs. However, current vision-centric paradigms often overlook the critical force and geometric feedback essential for complex tasks. This paper presents DeMUSE, a Deep Multimodal Unified Sparse Experts framework leveraging a Diffusion Transformer to integrate RGB, depth, and 6-axis force into a unified serialized stream. Adaptive Modality-specific Normalization (AdaMN) is employed to recalibrate modality-aware features, mitigating representation imbalance and harmonizing the heterogeneous distributions of multi-sensory signals. To facilitate efficient scaling, the architecture utilizes a Sparse Mixture-of-Experts (MoE) with shared experts, increasing model capacity for physical priors while maintaining the low inference latency required for real-time control. A Joint denoising objective synchronously synthesizes environmental evolution and action sequences to ensure physical consistency. Achieving success rates of 83.2% and 72.5% in simulation and real-world trials, DeMUSE demonstrates state-of-the-art performance, validating the necessity of deep multi-sensory integration for complex physical interactions.

具身智能多模态融合机器人操作扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。