arXiv:2409.00120cs.CLcs.AI2024-09中稿 · oral presentation …被引 3

针对英韩混用语境,提出统一对比学习与增强的嵌入方法ConCSE。

ConCSE: Unified Contrastive Learning and Augmentation for Code-Switched Embeddings

  • 融合对比学习与数据增强,提升混用语句的语义表示能力。
  • 在英韩混用文本相似度任务上平均提升1.77%性能。
  • 适用于需要处理多语言混用场景的自然语言处理研究者。

本文研究英语与韩语在同一语句中交织的代码混用(CS)现象,指出现有等价约束理论对英韩混用的复杂性捕捉不全,因两语言语法差异显著。为此,构建了专用于英韩混用场景的Koglish数据集:首先创建了Koglish-GLUE数据集,展示混用数据在多种任务中的重要性;发现不同基础多语言模型在单语与混用数据上表现差异明显。由此推测,虽在单语中表现优异的SimCSE在混用场景下存在局限。进一步基于混用增强方法构建了Koglish-NLI数据集进行验证。在此基础上,提出统一的对比学习与增强方法ConCSE,强调混用句子的语义特性。实验表明,ConCSE在Koglish-STS任务上实现平均1.77%的性能提升。

原文摘要 · Abstract (English)

This paper examines the Code-Switching (CS) phenomenon where two languages intertwine within a single utterance. There exists a noticeable need for research on the CS between English and Korean. We highlight that the current Equivalence Constraint (EC) theory for CS in other languages may only partially capture English-Korean CS complexities due to the intrinsic grammatical differences between the languages. We introduce a novel Koglish dataset tailored for English-Korean CS scenarios to mitigate such challenges. First, we constructed the Koglish-GLUE dataset to demonstrate the importance and need for CS datasets in various tasks. We found the differential outcomes of various foundation multilingual language models when trained on a monolingual versus a CS dataset. Motivated by this, we hypothesized that SimCSE, which has shown strengths in monolingual sentence embedding, would have limitations in CS scenarios. We construct a novel Koglish-NLI (Natural Language Inference) dataset using a CS augmentation-based approach to verify this. From this CS-augmented dataset Koglish-NLI, we propose a unified contrastive learning and augmentation method for code-switched embeddings, ConCSE, highlighting the semantics of CS sentences. Experimental results validate the proposed ConCSE with an average performance enhancement of 1.77\% on the Koglish-STS(Semantic Textual Similarity) tasks.

代码混用对比学习多语言嵌入表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。