针对东南亚电商低资源语言,提升多语言嵌入的鲁棒性与效率。
Compass-Embedding v4: Robust Contrastive Learning for Multilingual E-commerce Embeddings
- 设计类感知掩码机制,减少训练中的错误负样本。
- 构建合成数据与跨语言翻译语料,增强低资源语言覆盖。
- 结合量化与模型融合,实现高吞吐推理且不丢质量。
随着全球电商向新兴市场快速扩展,低资源语言缺乏高质量语义表示已成为检索、推荐和搜索系统的瓶颈。本文提出Compass-Embedding v4,一种专为东南亚(SEA)电商场景优化的高效多语言嵌入框架,应对数据稀缺、噪声标注和生产部署约束三大挑战。首先,针对混合任务监督下的大批次对比学习引入系统性错误负样本问题,提出轻量级类感知掩码(CAM),改进InfoNCE目标,在不降低训练效率的前提下提升语义区分能力。其次,针对低资源东南亚语言数据覆盖有限且不均的问题,通过上下文驱动的合成数据生成、跨语言翻译和结构化电商数据构建多样化训练语料,支持鲁棒的多语言与领域特定学习。第三,为满足生产环境对高吞吐推理的需求,结合鲁棒的大批次训练与球面模型合并策略以缓解灾难性遗忘,并通过vLLM和FP8量化优化推理性能。在多语言基准与自有电商任务上的大量实验表明,Compass-Embedding v4在主要东南亚语言上达到领先水平,显著优于通用嵌入模型在领域特定检索与分类任务中的表现,同时在高资源语言上保持竞争力。
原文摘要 · Abstract (English)
As global e-commerce rapidly expands into emerging markets, the lack of high-quality semantic representations for low-resource languages has become a decisive bottleneck for retrieval, recommendation, and search systems. In this work, we present Compass-Embedding v4, a high-efficiency multilingual embedding framework specifically optimized for Southeast Asian (SEA) e-commerce scenarios, where data scarcity, noisy supervision, and strict production constraints jointly challenge representation learning. Compass-Embedding v4 addresses three core challenges. First, large-batch contrastive training under mixed task supervision introduces systematic false negatives that degrade semantic alignment. We propose Class-Aware Masking (CAM), a lightweight modification to the InfoNCE objective that suppresses invalid in-batch negatives and improves semantic discrimination without altering training efficiency. Second, low-resource SEA languages suffer from limited and uneven data coverage. We construct a diversified training corpus through context-grounded synthetic data generation, cross-lingual translation, and structured e-commerce data construction, enabling robust multilingual and domain-specific learning. Third, production deployment requires high-throughput inference while preserving embedding quality. We combine robustness-driven large-batch training with spherical model merging to mitigate catastrophic forgetting, and optimize inference via vLLM and FP8 quantization. Extensive evaluations across multilingual benchmarks and proprietary e-commerce tasks show that Compass-Embedding v4 achieves state-of-the-art performance on major SEA languages, significantly outperforming general-purpose embedding models in domain-specific retrieval and classification, while maintaining competitive performance on high-resource languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。