arXiv:2511.12249cs.CL2025-11被引 1

ViConBERT提升越南语多义词语义理解能力

ViConBERT: Context-Gloss Aligned Vietnamese Word Embedding for Polysemous and Sense-Aware Representations

  • 融合对比学习与释义蒸馏,增强上下文嵌入
  • 在多义词消歧任务上达F1=0.87,优于基线模型
  • 首个大规模越南语语义评估数据集,适合语言研究者

近年来,上下文感知词嵌入在词义消歧(WSD)和语义相似度等任务中取得显著进展,但主要集中于英语等高资源语言。越南语仍缺乏鲁棒的语义建模工具和评估资源。本文提出ViConBERT,一种结合对比学习(SimCLR)与释义蒸馏的越南语上下文嵌入框架,以更好捕捉词义。同时构建了首个大规模合成数据集ViConWSD,涵盖WSD与上下文相似性任务。实验表明,ViConBERT在WSD任务上达到F1=0.87,ViCon任务上平均精度(AP)为0.88,ViSim-400上皮尔逊相关系数为0.60,有效建模离散词义与连续语义关系。代码、模型与数据已开源。

原文摘要 · Abstract (English)

Recent advances in contextualized word embeddings have greatly improved semantic tasks such as Word Sense Disambiguation (WSD) and contextual similarity, but most progress has been limited to high-resource languages like English. Vietnamese, in contrast, still lacks robust models and evaluation resources for fine-grained semantic understanding. In this paper, we present ViConBERT, a novel framework for learning Vietnamese contextualized embeddings that integrates contrastive learning (SimCLR) and gloss-based distillation to better capture word meaning. We also introduce ViConWSD, the first large-scale synthetic dataset for evaluating semantic understanding in Vietnamese, covering both WSD and contextual similarity. Experimental results show that ViConBERT outperforms strong baselines on WSD (F1 = 0.87) and achieves competitive performance on ViCon (AP = 0.88) and ViSim-400 (Spearman's rho = 0.60), demonstrating its effectiveness in modeling both discrete senses and graded semantic relations. Our code, models, and data are available at https://github.com/tkhangg0910/ViConBERT

越南语词嵌入多义词对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。