通过专家分离与动态负样本策略,实现多任务统一视觉嵌入的高效训练。
TSEmbed: Unlocking Task Scaling in Universal Multimodal Embeddings
- 融合专家混合与低秩适配,显式解耦冲突任务目标。
- 在MMEB和工业数据集上达到当前最优性能。
- 适合需要多任务统一嵌入的AI系统研发人员。
尽管多模态大语言模型具备出色的推理能力,但其转化为通用嵌入模型时仍受任务冲突严重制约。为此,我们提出TSEmbed框架,通过将专家混合(MoE)与低秩适配(LoRA)结合,显式分离冲突的任务目标。同时引入专家感知负采样(EANS),利用专家路由分布作为语义相似性的内在代理,动态优先选择与查询共享专家激活模式的有信息量难负样本,显著增强模型判别力并优化嵌入边界。为保障训练稳定性,设计两阶段学习范式,先固化专家专属性再通过EANS优化表示。TSEmbed在大规模多模态嵌入基准(MMEB)及真实工业生产数据集上均取得最先进性能,为通用多模态嵌入的任务级扩展奠定基础。
原文摘要 · Abstract (English)
Despite the exceptional reasoning capabilities of Multimodal Large Language Models (MLLMs), their adaptation into universal embedding models is significantly impeded by task conflict. To address this, we propose TSEmbed, a universal multimodal embedding framework that synergizes Mixture-of-Experts (MoE) with Low-Rank Adaptation (LoRA) to explicitly disentangle conflicting task objectives. Moreover, we introduce Expert-Aware Negative Sampling (EANS), a novel strategy that leverages expert routing distributions as an intrinsic proxy for semantic similarity. By dynamically prioritizing informative hard negatives that share expert activation patterns with the query, EANS effectively sharpens the model's discriminative power and refines embedding boundaries. To ensure training stability, we further devise a two-stage learning paradigm that solidifies expert specialization before optimizing representations via EANS. TSEmbed achieves state-of-the-art performance on both the Massive Multimodal Embedding Benchmark (MMEB) and real-world industrial production datasets, laying a foundation for task-level scaling in universal multimodal embeddings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。