arXiv:2511.12449cs.CVcs.AI2025-11中稿 · the IEEE Conferenc…被引 9

MOON2.0提升电商多模态理解,解决模态失衡与数据噪声问题。

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding

  • 动态专家混合模型按输入模态组成自适应处理,缓解模态不平衡。
  • 双层级对齐机制强化商品内图文语义关联,零样本性能达新高。
  • 图文协同增强加动态过滤,适合电商多模态研究与应用者。

近期多模态大模型显著推进了电商产品理解,但仍面临三大挑战:(i) 模态混合训练引发的模态不平衡;(ii) 商品内部视觉与文本信息间内在对齐关系利用不足;(iii) 对电商多模态数据中噪声的处理能力有限。为此,我们提出MOON2.0,一种面向电商产品理解的动态模态平衡多模态表征学习框架。其包含:(1) 基于模态驱动的专家混合(MoE)模型,根据输入样本的模态组成自适应处理,实现多模态联合学习以缓解模态不平衡;(2) 双层级对齐方法,更充分挖掘单个商品内的语义对齐特性;(3) 基于MLLM的图文协同增强策略,结合文本丰富与视觉扩展,并引入动态样本过滤以提升训练数据质量。我们还发布了MBE2.0——一个协同增强的电商多模态表征基准,用于学习与评估。实验表明,MOON2.0在MBE2.0及多个公开数据集上达到领先的零样本性能。基于注意力的热力图可视化提供了其多模态对齐能力增强的定性证据。

原文摘要 · Abstract (English)

Recent Multimodal Large Language Models (MLLMs) have significantly advanced e-commerce product understanding. However, they still face three challenges: (i) the modality imbalance induced by modality mixed training; (ii) underutilization of the intrinsic alignment relationships among visual and textual information within a product; and (iii) limited handling of noise in e-commerce multimodal data. To address these, we propose MOON2.0, a dynamic modality-balanced MultimOdal representation learning framework for e-commerce prOduct uNderstanding. It comprises: (1) a Modality-driven Mixture-of-Experts (MoE) that adaptively processes input samples by their modality composition, enabling Multimodal Joint Learning to mitigate the modality imbalance; (2) a Dual-level Alignment method to better leverage semantic alignment properties inside individual products; and (3) an MLLM-based Image-text Co-augmentation strategy that integrates textual enrichment with visual expansion, coupled with Dynamic Sample Filtering to improve training data quality. We further release MBE2.0, a co-augmented Multimodal representation Benchmark for E-commerce representation learning and evaluation at https://huggingface.co/datasets/ZHNie/MBE2.0. Experiments show that MOON2.0 delivers state-of-the-art zero-shot performance on MBE2.0 and multiple public datasets. Furthermore, attention-based heatmap visualization provides qualitative evidence of improved multimodal alignment of MOON2.0.

多模态电商理解表征学习MoE

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。