自动搜索多模态融合最佳架构,提升模型性能。
MixMAS: A Framework for Sampling-Based Mixer Architecture Search for Multimodal Fusion and Learning
- 基于采样微基准测试,自动筛选最优混合架构。
- 在多模态任务中显著优于固定架构设计。
- 适合需要高效融合图像、文本等异构数据的研究者。
为多模态数据融合选择合适的深度学习架构是一项挑战,因为需要有效整合具有不同结构和特性的多种数据类型。本文提出MixMAS,一种针对多模态学习的基于采样的混合适配架构搜索框架。该方法通过采样驱动的微基准测试策略,系统探索不同模态专用编码器、融合函数与融合网络的组合,自动识别满足特定任务性能指标的最优基于MLP的架构,适用于各类多模态机器学习(MML)任务。
原文摘要 · Abstract (English)
Choosing a suitable deep learning architecture for multimodal data fusion is a challenging task, as it requires the effective integration and processing of diverse data types, each with distinct structures and characteristics. In this paper, we introduce MixMAS, a novel framework for sampling-based mixer architecture search tailored to multimodal learning. Our approach automatically selects the optimal MLP-based architecture for a given multimodal machine learning (MML) task. Specifically, MixMAS utilizes a sampling-based micro-benchmarking strategy to explore various combinations of modality-specific encoders, fusion functions, and fusion networks, systematically identifying the architecture that best meets the task's performance metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。