用推理能力显式建模商品细粒度属性,提升电商产品理解效果。
MOON3.0: Reasoning-aware Multimodal Representation Learning for E-commerce Product Understanding
- 通过多头模态融合自适应整合原始输入信号
- 零样本下在多个任务上达到当前最优性能
- 适合需要精细商品分析的电商场景应用
随着电子商务的快速发展,探索通用表示而非任务特定表示受到越来越多关注。尽管近期多模态大语言模型(MLLMs)在产品理解方面取得显著进展,但通常仅作为特征提取器,隐式将产品信息编码为全局嵌入,限制了对细粒度属性的捕捉能力。因此,我们主张利用MLLM的推理能力显式建模细粒度商品属性,具有重要潜力。然而,实现这一目标面临三大挑战:(i) 长上下文推理易稀释模型对原始输入中关键信息的关注;(ii) 监督微调(SFT)主要鼓励机械模仿,限制有效推理策略的探索;(iii) 细粒度细节在前向传播过程中逐步衰减。为此,我们提出MOON3.0,首个面向产品表示学习的推理感知型MLLM模型。方法包括:(1) 采用多头模态融合模块自适应整合原始信号;(2) 引入联合对比与强化学习框架,自主探索更有效的推理策略;(3) 设计细粒度残差增强模块,逐层保留局部细节。此外,我们发布了大规模多模态电商基准MBE3.0。实验表明,该模型在自建基准及公开数据集上的多个下游任务中均实现了零样本下的先进性能。
原文摘要 · Abstract (English)
With the rapid growth of e-commerce, exploring general representations rather than task-specific ones has attracted increasing attention. Although recent multimodal large language models (MLLMs) have driven significant progress in product understanding, they are typically employed as feature extractors that implicitly encode product information into global embeddings, thereby limiting their ability to capture fine-grained attributes. Therefore, we argue that leveraging the reasoning capabilities of MLLMs to explicitly model fine-grained product attributes holds significant potential. Nevertheless, achieving this goal remains non-trivial due to several key challenges: (i) long-context reasoning tends to dilute the model's attention to salient information in the raw input; (ii) supervised fine-tuning (SFT) primarily encourages rigid imitation, limiting the exploration of effective reasoning strategies; and (iii) fine-grained details are progressively attenuated during forward propagation. To address these issues, we propose MOON3.0, the first reasoning-aware MLLM-based model for product representation learning. Our method (1) employs a multi-head modality fusion module to adaptively integrate raw signals; (2) incorporates a joint contrastive and reinforcement learning framework to autonomously explore more effective reasoning strategies; and (3) introduces a fine-grained residual enhancement module to progressively preserve local details throughout the network. Additionally, we release a large-scale multimodal e-commerce benchmark MBE3.0. Experimentally, our model demonstrates state-of-the-art zero-shot performance across various downstream tasks on both our benchmark and public datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。