arXiv:2510.16797cs.CLcs.AI2025-10Conference of the …

通过联合掩码与对比学习,提升文本嵌入模型在专业领域的适应能力。

MOSAIC: Masked Objective with Selective Adaptation for In-domain Contrastive Learning

  • 分阶段融合掩码语言建模与对比学习,实现领域自适应
  • 在高/低资源领域上,NDCG@10提升最高达13.4%
  • 适合需要精准领域表示的文本匹配与检索任务

我们提出MOSAIC(带选择性适配的掩码目标域适应框架),一种用于文本嵌入模型领域适应的多阶段框架,结合了领域特定的掩码监督。该方法解决将大规模通用领域文本嵌入模型适配到专业领域的问题。通过在统一训练流程中联合优化掩码语言建模(MLM)和对比学习目标,本方法在保留原始模型鲁棒语义区分能力的同时,有效学习领域相关表示。我们在高资源与低资源领域上进行了实证验证,相较于强基线模型,NDCG@10性能提升最高达13.4%。全面的消融实验进一步证明各组件的有效性,凸显均衡联合监督与分阶段适配的重要性。

原文摘要 · Abstract (English)

We introduce MOSAIC (Masked Objective with Selective Adaptation for In-domain Contrastive learning), a multi-stage framework for domain adaptation of text embedding models that incorporates joint domain-specific masked supervision. Our approach addresses the challenges of adapting large-scale general-domain text embedding models to specialized domains. By jointly optimizing masked language modeling (MLM) and contrastive objectives within a unified training pipeline, our method enables effective learning of domain-relevant representations while preserving the robust semantic discrimination properties of the original model. We empirically validate our approach on both high-resource and low-resource domains, achieving improvements up to 13.4% in NDCG@10 (Normalized Discounted Cumulative Gain) over strong general-domain baselines. Comprehensive ablation studies further demonstrate the effectiveness of each component, highlighting the importance of balanced joint supervision and staged adaptation.

领域适应对比学习掩码建模文本嵌入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。