用LaBSE和渐进式学习提升低资源语言的跨文化极化检测效果
Leveraging LaBSE with Progressive Curriculum Learning for Multicultural Polarization

- 将LaBSE用于极化检测,实现跨语言知识迁移
- 在低资源语言上最高提升0.2宏F1分数
- 适合关注多语种社会计算的研究者
在线极化检测在多语言多文化背景下仍具挑战性,尤其在低资源语言中因数据稀缺而困难。本文针对SemEval-2026任务9——多语言多文化在线极化检测,提出一种新架构:利用通常用于检索任务的LaBSE嵌入,实现强跨语言学习,显著提升低资源语言的性能,最高达0.2宏F1提升。同时,在基于检索的提示框架下,对Qwen系列多种编码器模型进行了全面消融实验。代码即将开源于https://github.com/carrycurious/PolarMind。
原文摘要 · Abstract (English)
Detecting online polarization remains a critical challenge, particularly in multilingual and multicultural contexts where intergroup hostility is prevalent. The problem is particularly challenging due to the data scarcity for these tasks in the low-resource languages. Identifying such phenomena has become an active area of research and is addressed in SemEval-2026 Task 9: Multilingual, Multicultural Online Polarization Detection. To address this problem we propose an architecture that leverages LaBSE embeddings - an unconventional choice typically reserved for retrieval tasks, to obtain strong cross-lingual learning which enhances scores in low-resource language by a score up to 0.2 macro F1. Furthermore, we provide a comprehensive ablation study evaluating the performance of diverse encoder models in the Qwen model family within a retrieval-based prompting framework. Our code will be soon available at https://github.com/carrycurious/PolarMind.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。