arXiv:2510.09293cs.CL2025-10Conference of the …被引 2

为句子同时生成显性与隐性语义向量,提升文本理解能力

One Sentence, Two Embeddings: Contrastive Learning of Explicit and Implicit Semantic Representations

  • 为每句分配两个嵌入:显性语义和隐性语义
  • 在信息检索与文本分类任务中性能显著提升
  • 适用于需要区分字面与深层含义的场景

句子嵌入方法虽取得显著进展,但仍难以捕捉句子中的隐性语义。这源于传统方法对每个句子仅赋予单一向量的固有局限。为此,我们提出DualCSE,一种为每句分配两个嵌入的句子嵌入方法:一个表征显性语义,另一个表征隐性语义。这两个嵌入共存于共享空间,可根据具体任务需求选择所需语义,如信息检索与文本分类。实验结果表明,DualCSE能有效编码显性与隐性语义,并提升下游任务性能。

原文摘要 · Abstract (English)

Sentence embedding methods have made remarkable progress, yet they still struggle to capture the implicit semantics within sentences. This can be attributed to the inherent limitations of conventional sentence embedding methods that assign only a single vector per sentence. To overcome this limitation, we propose DualCSE, a sentence embedding method that assigns two embeddings to each sentence: one representing the explicit semantics and the other representing the implicit semantics. These embeddings coexist in the shared space, enabling the selection of the desired semantics for specific purposes such as information retrieval and text classification. Experimental results demonstrate that DualCSE can effectively encode both explicit and implicit meanings and improve the performance of the downstream task.

句子嵌入对比学习隐性语义

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。