arXiv:2411.12156cs.CLcs.AI2024-11被引 1

用难负样本提升句子嵌入,让模型更懂语义

HNCSE: Advancing Sentence Embeddings via Hybrid Contrastive Learning with Hard Negatives

  • 引入难负样本增强对比学习,优化正负样本表示
  • 在语义相似度任务上超越现有方法,提升显著
  • 适合需要精准语义理解的NLP应用

无监督句子表示学习仍是现代自然语言处理中的关键挑战。近期,对比学习技术在捕捉文本语义方面取得显著进展。许多方法侧重于负样本的优化。在计算机视觉领域,难负样本(接近决策边界、难以区分的样本)已被证明能增强表示学习。然而,由于文本的句法和语义结构复杂,将难负样本应用于句子对比学习仍具挑战。为此,我们提出HNCSE,一种基于领先方法SimCSE的新型对比学习框架。其核心在于创新性地利用难负样本,同时增强正负样本的学习效果,实现更深层的语义理解。在语义文本相似度及迁移任务数据集上的实证测试验证了HNCSE的优越性。

原文摘要 · Abstract (English)

Unsupervised sentence representation learning remains a critical challenge in modern natural language processing (NLP) research. Recently, contrastive learning techniques have achieved significant success in addressing this issue by effectively capturing textual semantics. Many such approaches prioritize the optimization using negative samples. In fields such as computer vision, hard negative samples (samples that are close to the decision boundary and thus more difficult to distinguish) have been shown to enhance representation learning. However, adapting hard negatives to contrastive sentence learning is complex due to the intricate syntactic and semantic details of text. To address this problem, we propose HNCSE, a novel contrastive learning framework that extends the leading SimCSE approach. The hallmark of HNCSE is its innovative use of hard negative samples to enhance the learning of both positive and negative samples, thereby achieving a deeper semantic understanding. Empirical tests on semantic textual similarity and transfer task datasets validate the superiority of HNCSE.

句子嵌入对比学习难负样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。