融合视觉触觉与多智能体关系图,动态调整感知权重提升协作抓取稳定性。
ADM-DP: Adaptive Dynamic Modality Diffusion Policy through Vision-Tactile-Graph Fusion for Multi-Agent Manipulation
- 通过跨模态调制增强视觉感知,融合图像与点云信息。
- 触觉反馈驱动抓取修正,使抓握成功率提升12%-25%。
- 自适应注意力机制动态分配模态权重,适合复杂协作任务。
多智能体机器人操作因协调、抓握稳定性和共享空间内避障的综合需求而面临挑战。为此,我们提出自适应动态模态扩散策略(ADM-DP),通过融合视觉、触觉和基于图结构的多智能体位姿信息实现协同控制。该框架包含四项创新:首先,采用特征逐通道线性调制(FiLM)融合RGB与点云特征以增强感知;其次,利用力敏电阻(FSR)反馈检测接触不足并触发抓取修正,提升抓握稳定性;第三,基于图的碰撞编码器将多个智能体的工具中心点(TCP)位置作为结构化运动学上下文,维持空间感知并减少相互干扰;第四,自适应模态注意力机制(AMAM)根据任务上下文动态重加权各模态输入,实现灵活融合。为保证可扩展性与模块化,采用解耦训练范式,各智能体独立学习策略但共享空间信息,既降低依赖性又保持集体意识。在七项多智能体任务中,ADM-DP相较最先进基线性能提升12%-25%。消融实验表明,在需多感官融合的任务中改进最显著,验证了自适应融合策略的有效性与鲁棒性。
原文摘要 · Abstract (English)
Multi-agent robotic manipulation remains challenging due to the combined demands of coordination, grasp stability, and collision avoidance in shared workspaces. To address these challenges, we propose the Adaptive Dynamic Modality Diffusion Policy (ADM-DP), a framework that integrates vision, tactile, and graph-based (multi-agent pose) modalities for coordinated control. ADM-DP introduces four key innovations. First, an enhanced visual encoder merges RGB and point-cloud features via Feature-wise Linear Modulation (FiLM) modulation to enrich perception. Second, a tactile-guided grasping strategy uses Force-Sensitive Resistor (FSR) feedback to detect insufficient contact and trigger corrective grasp refinement, improving grasp stability. Third, a graph-based collision encoder leverages shared tool center point (TCP) positions of multiple agents as structured kinematic context to maintain spatial awareness and reduce inter-agent interference. Fourth, an Adaptive Modality Attention Mechanism (AMAM) dynamically re-weights modalities according to task context, enabling flexible fusion. For scalability and modularity, a decoupled training paradigm is employed in which agents learn independent policies while sharing spatial information. This maintains low interdependence between agents while retaining collective awareness. Across seven multi-agent tasks, ADM-DP achieves 12-25% performance gains over state-of-the-art baselines. Ablation studies show the greatest improvements in tasks requiring multiple sensory modalities, validating our adaptive fusion strategy and demonstrating its robustness for diverse manipulation scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。