用微调大模型检测多语言极化,提升跨文化语境下的识别鲁棒性。
mdok-style at SemEval-2026 Task 9: Finetuning LLMs for Multilingual Polarization Detection
- 采用QLoRA技术微调中等规模大模型进行序列分类。
- 通过数据增强使模型在22种语言上对极化信号更鲁棒。
- 适合关注多语言内容安全与社会情绪分析的研究者。
SemEval-2026 Task 9聚焦于多语言极化检测,涵盖多语言、多文化、多事件背景下沿三个维度(子任务)的极化识别、类型判断与表现形式分析。在线极化现象令人担忧,因其常伴随仇恨言论、攻击性话语与社会分裂。提前检测极化对构建更安全包容的网络空间至关重要。我们通过使用QLoRA参数高效微调技术,对中等规模大模型进行微调,以完成序列分类任务。训练数据通过引入匿名化、小写、大写及同形异义字符变体,增强了22种语言的训练集,从而提升模型在多语言场景下的检测鲁棒性。
原文摘要 · Abstract (English)
SemEval-2026 Task 9 is focused on multilingual polarization detection. Specifically, it covers the identification of multilingual, multicultural and multievent polarization along three axes (in subtasks), namely detection, type, and manifestation. Online polarization presents a concern, because it is often followed by hate speech, offensive discourse, and social fragmentation. Therefore, its detection before it escalates is crucial for a safer and more inclusive online space. We have coped with this SemEval task by finetuning mid-size LLMs for the sequence-classification task using the QLoRA parameter-efficient finetuning technique. The training data augmented the multilingual (22 languages) training sets by anonymized, lower-cased, upper-cased, and homoglyphied counterparts, making the detection more robust.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。