提出HMGCLIP框架,提升电商产品细粒度表征能力
HMGCLIP: Heterogeneous Multi-Granularity Contrastive Learning for E-commerce Representation Learning

- 构建异构超图挖掘结构感知难负例,实现多粒度语义对齐
- 在新数据集和MAVE上显著优于现有模型,细粒度任务提升12.3%
- 适合需要精准属性识别的电商推荐与检索场景
尽管近期多模态大模型(MLLMs)在通用产品理解上取得进展,但其隐式将产品信息编码为全局嵌入,限制了对细粒度属性的捕捉能力,影响在需精确属性区分任务(如视觉相似产品间的材质差异)中的表现。为此,我们提出HMGCLIP,一种统一的多模态嵌入框架。通过构建异构超图,利用超图拓扑挖掘结构感知的难负例,并在关系与超边层面对齐多粒度语义。该设计支持双粒度推理机制,动态融合属性证据以适应细粒度与粗粒度下游任务。此外,我们发布了一个全面的细粒度电商数据集,以促进未来基准测试。在该新数据集及公开的MAVE基准上的大量实验表明,HMGCLIP超越强多模态编码器、MLLMs及电商基线,验证了其优越性。
原文摘要 · Abstract (English)
Although recent Multimodal Large Language Models (MLLMs) have advanced general product understanding, they implicitly encode product information into global embeddings, thereby limiting their ability to capture fine-grained attributes. This limitation hinders performance in tasks requiring precise attribute discrimination, such as distinguishing subtle material differences among visually similar products. To address this challenge, we propose HMGCLIP, a unified multimodal embedding framework. By constructing a heterogeneous hypergraph, we leverage hypergraph topology to mine structure-aware hard negatives and align multi-granular semantics at both relation and hyperedge levels. This design enables a dual-granularity inference mechanism that dynamically fuses attribute evidence for both fine-grained and coarse-grained downstream tasks. Furthermore, we release a comprehensive fine-grained e-commerce dataset to facilitate future benchmarking. Extensive experiments on this new dataset and the public MAVE benchmark show that HMGCLIP outperforms strong multimodal encoders, MLLMs, and e-commerce baselines, validating the superiority of HMGCLIP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。