arXiv:2412.16243cs.LG2024-12被引 9

提出多模态AutoML实用技巧,提升图像文本表格数据融合效果

Bag of Tricks for Multimodal AutoML with Image, Text, and Tabular Data

  • 系统测试多种多模态融合与数据增强策略
  • 在22个真实数据集上实现稳定高精度表现
  • 适合需要处理多源异构数据的开发者参考

本文研究自动机器学习(AutoML)的最佳实践。尽管以往工作主要聚焦于单模态数据,多模态场景仍鲜有探索。本研究针对包含图像、文本和表格数据的灵活组合的分类与回归问题,构建了一个涵盖22个来自真实应用的多模态数据集的基准,覆盖三种模态的所有四种组合。在该基准上,我们深入分析了多模态融合策略、多模态数据增强、将表格数据转为文本、跨模态对齐以及缺失模态处理等设计选择。通过大量实验与分析,提炼出一系列有效策略,并整合为统一流水线,在多样化数据集上实现了稳健性能。

原文摘要 · Abstract (English)

This paper studies the best practices for automatic machine learning (AutoML). While previous AutoML efforts have predominantly focused on unimodal data, the multimodal aspect remains under-explored. Our study delves into classification and regression problems involving flexible combinations of image, text, and tabular data. We curate a benchmark comprising 22 multimodal datasets from diverse real-world applications, encompassing all 4 combinations of the 3 modalities. Across this benchmark, we scrutinize design choices related to multimodal fusion strategies, multimodal data augmentation, converting tabular data into text, cross-modal alignment, and handling missing modalities. Through extensive experimentation and analysis, we distill a collection of effective strategies and consolidate them into a unified pipeline, achieving robust performance on diverse datasets.

多模态AutoML数据融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。