arXiv:2608.04348cs.CVcs.AI2026-08中稿 · presentation at th…

通过结构化特征排序提升图像与表格数据的多模态学习效果

iStructTab: Structured Feature Sequencing for Multimodal Learning of Image and Tabular Data

论文配图:iStructTab: Structured Feature Sequencing for Multimodal Learning of Image and Tabular Data
图 1 · 摘自论文原文
  • 基于相似性图计算优化特征描述符,实现有序特征序列生成
  • 在多个基准上显著降低特征分散性,提升预测性能与鲁棒性
  • 适合关注多模态特征对齐与结构化建模的研究者

图像与表格数据的多模态学习常因表示无效而面临冗余、分散及泛化问题。为此,我们提出图增强描述符排序(GEDS),基于列排列问题(CPP)原理,通过相似性图计算优化特征统计描述符,系统确定有效特征序列。将GEDS嵌入顺序感知的高效Transformer框架中,利用顺序感知记忆标记,通过专用损失函数显式遵循生成的特征序列。在多个多模态基准上的实验表明,iStructTab能有效减少特征分散,提升预测性能与鲁棒性,凸显结构化特征排序在多模态学习中的重要性。

原文摘要 · Abstract (English)

Multimodal learning of images and tabular data is often impaired by ineffective representations, resulting in redundancy, dispersion, and generalization problems. To tackle this challenge, we introduce Graph-Enhanced Descriptor Sequencing (GEDS), a structured feature sequencing algorithm grounded in principles from the Column Permutation Problem (CPP). GEDS refines statistical descriptors of the features through similarity graph-based computations, systematically determining an effective feature sequencing. We incorporate GEDS within an order-aware efficient transformer framework, utilizing order-aware memory tokens that explicitly adhere to the derived feature sequencing via a dedicated loss function. Experimental results across multimodal benchmarks demonstrate that iStructTab effectively minimizes feature dispersion, improving predictive performance and robustness, and highlighting the significance of structured feature sequencing in multimodal learning.

多模态学习特征排序图像表格融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。