arXiv:2409.12737cs.CLcs.AI2024-09ACL被引 5

同时优化句子和词语级别目标,提升跨语言句子表征质量

MEXMA: Token-level objectives improve sentence representations

  • 引入句子与词粒度双重目标联合训练
  • 在多任务上显著提升句子表示性能
  • 适合需要精准语义对齐的跨语言应用

当前预训练跨语言句子编码器仅使用句子级目标,可能导致词级别信息丢失,从而影响句子表征质量。本文提出MEXMA,一种融合句子级与词级目标的新方法:利用一种语言的句子表征预测另一种语言中的掩码词,且句子表征与所有词均直接更新编码器。实验表明,加入词级目标可显著提升多个任务上的句子表征质量。该方法在双语挖掘及多个下游任务中优于现有预训练跨语言句子编码器。我们还分析了词中编码的信息及其如何构建句子表征。

原文摘要 · Abstract (English)

Current pre-trained cross-lingual sentence encoders approaches use sentence-level objectives only. This can lead to loss of information, especially for tokens, which then degrades the sentence representation. We propose MEXMA, a novel approach that integrates both sentence-level and token-level objectives. The sentence representation in one language is used to predict masked tokens in another language, with both the sentence representation and all tokens directly updating the encoder. We show that adding token-level objectives greatly improves the sentence representation quality across several tasks. Our approach outperforms current pre-trained cross-lingual sentence encoders on bi-text mining as well as several downstream tasks. We also analyse the information encoded in our tokens, and how the sentence representation is built from them.

跨语言句子表征词粒度预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。