arXiv:2607.10312cs.CLcs.AI2026-07

跨语言极化检测:用混合模型适配英豪两种语言

Polarization Detection: A Hybrid Approach with AfroXLMR-Social and DeBERTa for Low- and High-Resource Settings

  • 英语用DeBERTa,豪萨语用AfroXLMR-Social,针对性选型
  • 在低资源豪萨语上表现稳健,多任务准确率超基准
  • 结合LoRA与数据增强,适合资源少的极化分析场景

在线极化现象迅速蔓延,威胁社会凝聚力,亟需在多种语言环境下高效运行的自动化检测系统。本文介绍我们参与2026年POLAR共享任务的系统方案,聚焦英语和豪萨语中极化话语的检测与特征刻画。提出一种混合建模策略:英语二分类任务采用单语强项的DeBERTa;豪萨语及所有细粒度子任务(类型与表现形式)则使用领域适配的多语言模型AfroXLMR-Social,该模型在捕捉社交媒体文本中的极化细微差异方面表现关键。为应对计算限制与数据稀缺,引入低秩微调(LoRA)及通过nlpaug进行文本数据增强。在所有三个子任务中均取得有竞争力的结果,证明根据子任务需求定制模型选择,能实现性能最优平衡。

原文摘要 · Abstract (English)

The rapid proliferation of online polarization threatens social cohesion, necessitating robust automated detection systems that operate effectively across diverse linguistic contexts. This paper presents our system description for the POLAR Shared Task 2026, focusing on the detection and characterization of polarized discourse in English and Hausa. We propose a hybrid modeling strategy: for English binary detection, we leverage the monolingual strength of \textbf{DeBERTa}, while for Hausa and all fine-grained subtasks (Types and Manifestations), we utilize \textbf{AfroXLMR-Social}. This domain-adapted multilingual model proved critical for capturing the nuances of polarization in social media text. To further address computational constraints and data scarcity, we implement Low-Rank Adaptation (LoRA) and textual data augmentation via \texttt{nlpaug}. We report competitive results across all three subtasks, demonstrating that model selection tailored to specific subtask requirements yields the best balance of performance.

极化检测多语言低资源混合模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。