arXiv:2507.13364cs.CVcs.AI2025-07CVPR被引 50

统一12种模态的多任务学习模型,支持跨模态联合训练。

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning

  • 用专用分词器+共享Transformer+交叉注意力融合多模态数据
  • 在25个数据集上达到顶尖性能,覆盖12种模态
  • 适合需要多模态融合与多任务学习的研究者

我们提出一种新型多模态多任务网络及其训练算法,可处理约12种不同模态的数据:图像、视频、音频、文本、深度图、点云、时间序列、表格数据、图结构、X光、红外、惯性测量单元(IMU)和高光谱数据。该方法采用模态专用分词器、共享Transformer架构和交叉注意力机制,将不同模态数据映射到统一嵌入空间。通过为各模态特定任务配置模态专属任务头,实现多模态与多任务场景下的联合建模。提出一种基于迭代模态切换的预训练策略,并设计了一种训练算法,在全模态联合训练与成对模态训练之间进行权衡。在来自12种模态的25个数据集上进行了全面评估,结果表明该架构、预训练策略及适配的多任务训练方法均具有显著有效性。

原文摘要 · Abstract (English)

We present a novel multimodal multitask network and associated training algorithm. The method is capable of ingesting data from approximately 12 different modalities namely image, video, audio, text, depth, point cloud, time series, tabular, graph, X-ray, infrared, IMU, and hyperspectral. The proposed approach utilizes modality specialized tokenizers, a shared transformer architecture, and cross-attention mechanisms to project the data from different modalities into a unified embedding space. It addresses multimodal and multitask scenarios by incorporating modality-specific task heads for different tasks in respective modalities. We propose a novel pretraining strategy with iterative modality switching to initialize the network, and a training algorithm which trades off fully joint training over all modalities, with training on pairs of modalities at a time. We provide comprehensive evaluation across 25 datasets from 12 modalities and show state of the art performances, demonstrating the effectiveness of the proposed architecture, pretraining strategy and adapted multitask training.

多模态Transformer多任务学习统一建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。