arXiv:2411.18990cs.CLcs.AI2024-11被引 2

通过降维白化减少语言偏置,提升跨语言文本相关性判断精度

USTCCTSU at SemEval-2024 Task 1: Reducing Anisotropy for Cross-lingual Semantic Textual Relatedness Task

  • 基于XLM-R-base模型,采用白化技术降低语义表示的各向异性
  • 在西班牙语任务中获第2名,印尼语任务中获第3名,多语言任务进入前十
  • 提出数据过滤方法缓解多语言困境,适合跨语言理解研究者参考

跨语言语义文本相关性任务是解决跨语言沟通与文本理解挑战的重要研究方向,有助于建立不同语言间的语义关联,对机器翻译、多语言信息检索和跨语言文本理解等下游任务至关重要。基于大量对比实验,我们选择XLM-R-base作为基础模型,并采用基于白化的预训练句子表示以降低各向异性。此外,针对给定训练数据,设计了精细的数据过滤方法以缓解多语言困境。通过该方法,在西班牙语任务中取得第2名,在印尼语任务中取得第3名,多个语言任务进入竞赛第10名以内。我们还进行了全面分析,旨在为未来提升跨语言任务表现提供启发。

原文摘要 · Abstract (English)

Cross-lingual semantic textual relatedness task is an important research task that addresses challenges in cross-lingual communication and text understanding. It helps establish semantic connections between different languages, crucial for downstream tasks like machine translation, multilingual information retrieval, and cross-lingual text understanding.Based on extensive comparative experiments, we choose the XLM-R-base as our base model and use pre-trained sentence representations based on whitening to reduce anisotropy.Additionally, for the given training data, we design a delicate data filtering method to alleviate the curse of multilingualism. With our approach, we achieve a 2nd score in Spanish, a 3rd in Indonesian, and multiple entries in the top ten results in the competition's track C. We further do a comprehensive analysis to inspire future research aimed at improving performance on cross-lingual tasks.

跨语言语义相关性模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。