解决多模态触觉学习中的单域泛化难题,提升跨域适应能力。
OmniVaT: Single Domain Generalization for Multimodal Visual-Tactile Learning
- 通过分数傅里叶适配器统一视觉与触觉嵌入空间,缓解模态差异。
- 在无多域训练数据下,跨域泛化性能超越现有方法。
- 适合需要强泛化能力的机器人感知系统开发者。
视觉-触觉学习(VTL)使具身智能体通过融合视觉(VIS)和触觉(TAC)传感器感知物理世界。然而,VTL仍面临视觉与触觉图像间的模态差异,以及因非标准化触觉传感器和不一致数据采集流程导致的域差距问题。本文将这些挑战定义为一项新任务:多模态视觉-触觉学习中的单域泛化(SDG-VTL)。我们提出OmniVaT框架,首次成功应对该任务。一方面,OmniVaT引入多模态分数傅里叶适配器(MFFA),将VIS和TAC嵌入映射至统一的嵌入-频率空间,有效缓解模态差距,无需多域训练数据或精细的跨模态融合策略。另一方面,其还包含离散树生成(DTG)模块,通过分层树结构生成多样且可靠的多模态分数表示,增强对未见域中波动性域偏移的适应性。大量实验表明,OmniVaT在SDG-VTL任务上表现出卓越的跨域泛化性能。
原文摘要 · Abstract (English)
Visual-tactile learning (VTL) enables embodied agents to perceive the physical world by integrating visual (VIS) and tactile (TAC) sensors. However, VTL still suffers from modality discrepancies between VIS and TAC images, as well as domain gaps caused by non-standardized tactile sensors and inconsistent data collection procedures. We formulate these challenges as a new task, termed single domain generalization for multimodal VTL (SDG-VTL). In this paper, we propose an OmniVaT framework that, for the first time, successfully addresses this task. On the one hand, OmniVaT integrates a multimodal fractional Fourier adapter (MFFA) to map VIS and TAC embeddings into a unified embedding-frequency space, thereby effectively mitigating the modality gap without multi-domain training data or careful cross-modal fusion strategies. On the other hand, it also incorporates a discrete tree generation (DTG) module that obtains diverse and reliable multimodal fractional representations through a hierarchical tree structure, thereby enhancing its adaptivity to fluctuating domain shifts in unseen domains. Extensive experiments demonstrate the superior cross-domain generalization performance of OmniVaT on the SDG-VTL task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。