改进多语言句子嵌入模型,统一处理词语上下文的序数与二分类任务。
XL-DURel: Finetuning Sentence Transformers for Ordinal Word-in-Context Classification
- 基于复数空间角度距离设计排序损失,提升序数分类效果。
- 在序数与二分类任务上均超越现有模型性能。
- 证明二分类可视为序数分类特例,适合统一建模场景。
我们提出 XL-DURel,一种针对序数词语上下文(Word-in-Context, WiC)分类优化的微调多语言句子嵌入模型。通过测试多种回归与排序任务的损失函数,我们发现基于复数空间角度距离的排序目标,在序数和二分类数据上均优于先前模型。进一步表明,二分类 WiC 可视为序数 WiC 的特例;对通用序数任务进行优化,能提升在特定二分类任务上的表现。这为不同任务形式下的 WiC 建模提供了统一框架。
原文摘要 · Abstract (English)
We propose XL-DURel, a finetuned, multilingual Sentence Transformer model optimized for ordinal Word-in-Context classification. We test several loss functions for regression and ranking tasks managing to outperform previous models on ordinal and binary data with a ranking objective based on angular distance in complex space. We further show that binary WiC can be treated as a special case of ordinal WiC and that optimizing models for the general ordinal task improves performance on the more specific binary task. This paves the way for a unified treatment of WiC modeling across different task formulations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。