通过分层注意力增强多模态分类鲁棒性,有效应对噪声与长尾问题。
FLUID: Flow-Latent Unified Integration via Token Distillation for Expert Specialization in Multimodal Learning
- 用可学习查询令牌提取跨模态关键特征,实现细粒度信息融合。
- 在GLAMI-1M上达91%准确率,显著优于基线并抗标签噪声与类别不均衡。
- 轻量级专家混合机制支持高效语义专精,适合复杂商品分类任务。
多模态分类需稳健整合视觉与文本信号,但常见融合策略易受模态特异性噪声影响。本文提出 extsc{FLUID}——基于令牌蒸馏的流-潜在统一融合框架,用于多模态学习中的专家专精。其核心包括:(1) 可学习查询令牌(Q-transforms),从模态专用主干中提炼并保留关键令牌级特征;(2) 两阶段融合:先通过对比对齐强化跨模态一致性,再经门控机制与Q瓶颈实现任务感知的自适应融合与信息压缩;(3) 预测时轻量级、负载均衡的专家混合机制,支持对多样语义模式的高效专精。大量实验表明, extsc{FLUID}在GLAMI-1M基准上达到91%准确率,显著超越已有基线,在标签噪声、长尾类别失衡和语义异质性下仍具强鲁棒性。消融研究验证了各组件的独立与协同优势,证明 extsc{FLUID}是可扩展、抗噪的多模态产品分类解决方案。
原文摘要 · Abstract (English)
Multimodal classification requires robust integration of visual and textual signals, yet common fusion strategies are brittle and vulnerable to modality-specific noise. In this paper, we present \textsc{FLUID}-Flow-Latent Unified Integration via Token Distillation for Expert Specialization, a principled token-level pipeline that improves cross-modal robustness and scalability. \textsc{FLUID} contributes three core elements: (1) \emph{Q-transforms}, learnable query tokens that distill and retain salient token-level features from modality-specific backbones; (2) a two-stage fusion scheme that enforces cross-modal consistency via contrastive alignment and then performs adaptive, task-aware fusion through a gating mechanism and a \emph{Q-bottleneck} that selectively compresses information for downstream reasoning; and (3) a lightweight, load-balanced Mixture-of-Experts at prediction time that enables efficient specialization to diverse semantic patterns. Extensive experiments demonstrate that \textsc{FLUID} attains \(91\%\) accuracy on the GLAMI-1M benchmark, significantly outperforming prior baselines and exhibiting strong resilience to label noise, long-tail class imbalance, and semantic heterogeneity. Targeted ablation studies corroborate both the individual and synergistic benefits of the proposed components, positioning \textsc{FLUID} as a scalable, noise-resilient solution for multimodal product classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。