arXiv:2607.15861cs.CLcs.AI2026-07中稿 · ETTIS 2026

针对多语言混杂文本,提出条件化可信度融合方法提升毒性检测效果。

Conditional Reliability of Toxicity Signals for Multilingual and Code-Mixed Abuse Detection

  • 基于编码器表示动态调整外部毒性信号权重,实现条件化融合。
  • 在12个本地场景中10次、8个跨域场景中7次优于基线模型。
  • 对敏感内容如辱骂、暴力威胁等高风险片段提升显著,适合安全审核场景。

内容审核系统越来越多依赖外部毒性检测工具,但在多语言混杂、音译、俚语和语言错配情况下这些工具可靠性下降。本文研究印度多语言及混杂短文本中毒性先验的条件可靠性:英语毒性、印地语类攻击性内容和规则级严重性线索可作为有用证据,但仅在特定语言与严重性情境下有效。我们提出ToxGate,一种信任融合头,将每个辅助信号根据编码器表示进行条件化处理后再融入预测状态。在三个短文本滥用数据集、四种Transformer编码器及每组五次随机种子下,ToxGate在12个本地设置中10次、8个迁移设置中7次超越匹配的纯编码器模型。最大且最可解释的增益出现在高风险审核子集,包括明确辱骂、暴力威胁及跨数据集迁移场景。核心启示是:审核系统应将外部毒性工具与先验视为条件证据而非固定特征;在细分消融实验中,源特定门控在迁移、严重滥用片段和高风险筛选中表现最优。

原文摘要 · Abstract (English)

Moderation systems increasingly rely on external toxicity tools, but those tools are unreliable under code-mixing, transliteration, slang, and language mismatch. We study the \emph{conditional reliability} of toxicity priors in Indian multilingual and code-mixed short text: English toxicity, Indic abuse, and rule-based severity cues can be useful evidence, but only in some linguistic and abuse-severity contexts. We propose ToxGate, a trust-fusion head that conditions each auxiliary signal on the encoder representation before adding it to the prediction state. Across three short-text abuse datasets, four transformer encoders, and five seeds per setting, ToxGate improves over matched plain encoders in 10 of 12 in-domain settings and 7 of 8 transfer settings. The largest and most interpretable gains occur in high-risk moderation slices, including explicit slurs, violent threats, and cross-dataset transfer. The broader lesson is that moderation systems should treat external toxicity tools and priors as conditional evidence rather than fixed features or ground truth, in focused ablations, source-specific gating gives the strongest results in transfer, severe-abuse slices, and high-risk triage.

内容审核多语言混合语言可信度融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。