arXiv:2607.23512cs.CLcs.SY2026-07

剖析辱骂语言检测跨域失效根源,提出可量化优化方案

The Cross-Domain Generalization Cost of Offensive Language Detection

  • 将性能下降分解为数据集与语言双重影响,独立测量其作用
  • 发现数据集差异导致的损失远超语言差异,是主因
  • 提出联合训练策略,实现多语言能力提升与原任务保全的可控权衡

辱骂语言检测模型在跨数据集和跨语言场景下普遍存在性能下降问题,但现有研究多停留在现象描述,缺乏系统性方法对退化原因进行拆解并量化修复成本。本文提出一个诊断与优化框架,包含三个协同技术组件:首先,通过零样本迁移损失分解,将从 OLID 到 MLMA 的性能下降分离为可独立测量的数据集效应与语言效应;其次,设计受控微调协议,通过比较持续微调与冷启动下的少样本学习曲线,量化适应效率及对源任务的隐性损伤;第三,提出三种融合温度采样与经验回放的联合训练策略,实现多语言能力提升与源任务性能保持之间的可控帕累托权衡。实验表明,数据集效应主导零样本迁移损失,显著超过语言效应。无回放机制的少样本适配虽数据高效,但对源任务造成的损伤是联合训练策略的4至9倍,且损伤极不稳定。三种联合训练策略以牺牲3.2至4.1个百分点的源任务性能,换取8.1至42.6个百分点的多语言能力增益,形成清晰可控的帕累托边界。

原文摘要 · Abstract (English)

Offensive language detection models generally suffer performance degradation when deployed across datasets and across languages, yet most existing studies stop at reporting this phenomenon and lack a systematic methodology for decomposing the causes of degradation into attributable components and quantifying the cost of remediation. This paper proposes a diagnosis and optimization framework composed of three coordinated technical components. First, a zero-shot transfer loss decomposition that separates the performance degradation from OLID to MLMA into two independently measurable components, namely dataset effect and language effect. Second, a controlled fine-tuning protocol that quantifies both adaptation efficiency and the hidden damage inflicted on the source task by comparing few shot learning curves under continued fine-tuning and cold-start starting points. Third, three joint training strategies incorpo rating temperature sampling and experience replay, which offer a controllable Pareto trade-off between improving multilingual capability and preserving source-task performance. Experiments built on this framework show that the dataset effect dominates the zero-shot transfer loss and substantially outweighs the language effect. Few-shot adaptation without a replay mechanism, though data-efficient, inflicts source task damage 4 to 9 times greater than that of the joint training strategies, and its damage magnitude is highly unstable. The three joint training strategies trade 3.2 to 4.1 percentage points of source-task performance for 8.1 to 42.6 percentage points of multilingual capability gain, forming a clear and controllable Pareto trade-off.

自然语言处理跨域泛化模型鲁棒性多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。