代码混杂内容导致仇恨内容审核决策不稳定,影响真实系统运行。
When Surface Form Changes Moderation Decisions: A Paired Study of Code-Mixed Workflow Instability

- 对比纯英文与泰米尔-英文混用内容的审核决策差异
- 混用输入使误判率升至10.4%,需审阅比例翻倍至29.7%
- 适合关注现实审核系统鲁棒性的研究者阅读
仇恨内容审核常以纯净英文输入为评估标准,但实际系统需对内容做出允许(ALLOW)、标记(FLAG)或审查(REVIEW)等操作。本文通过配对评估设置,比较同一内容在纯净英文与泰米尔-英文混用表达下的审核表现。在基于纯净英文数据调优的阈值下,混用输入导致显著的决策不稳定性,配对清洁-混用决策翻转率达0.265。主要影响包括审查负担增加(从0.138升至0.297)和非仇恨内容误标率上升(从0.069增至0.104)。泰米尔语独用输入表现更差,表明问题源于语言覆盖不足而非混用特有现象。简单基于分歧的延迟规则可降低自动错误,但需更高审查负荷。结果表明,工作流层面评估能揭示分类指标忽略的审核失效。
原文摘要 · Abstract (English)
Hate moderation is often evaluated as classification on clean English inputs, but deployed systems must route content to actions such as ALLOW, FLAG, or REVIEW. We study how this workflow changes under code-mixed inputs using a paired evaluation setting where the same underlying content is expressed as clean English and Tamil-English code-mix. Under thresholds tuned on clean English development data, code-mixed inputs produce substantial action instability, with a paired clean- to-code-mix decision flip rate of 0.265. The main workflow effects are increased review burden and increased false-flagging of non-hateful content: review rate rises from 0.138 to 0.297 and non-hate false-flag rate rises from 0.069 to 0.104. Tamil-only inputs show stronger degradation overall, suggesting a broader language-coverage limitation rather than the same code-mixed instability pattern. A simple disagreement-based deferral rule reduces automatic errors on stressed inputs, but only by increasing review load. These results show that workflow-level evaluation reveals moderation failures that standard classification summaries can miss.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。