毒性检测应关注语境中的伤害,而非文本本身是否坏。
Toxicity Detection Should Measure Contextual Harm, Not Text-Intrinsic Badness
- 将毒性定义为语境中规范违背与情绪压力的关联关系
- 揭示传统方法误标方言、再利用语言及隐晦攻击
- 适合安全系统设计者和平台内容策略制定者
毒性检测已成为在线内容审核、数据集过滤及部署语言模型的核心安全机制。然而,大多数检测器仍将毒性视为孤立文本的固有属性。本文主张,毒性检测应作为对具体语境中沟通行为造成伤害的测量,而非单一标签的文本分类任务。毒性并非仅存在于词汇本身,而是在受众于特定规范与社会背景下解读交际行为时产生。我们提出情境压力框架(CSF),将毒性定义为感知规范违背与引发的压力或干扰之间的关系。该框架解释了为何基于文本固有的检测器会误判方言或被重新使用的语言,遗漏隐晦或语用性攻击,并在保持语义的前提下发生脆弱性。我们进一步提出CSF-Eval评估方案,区分文本风险、规范违背、干扰程度、不确定性及政策应对等维度。
原文摘要 · Abstract (English)
Toxicity detection has become core safety infrastructure for online moderation, dataset filtering, and deployed language-model systems. Yet most detectors still treat toxicity as an intrinsic property of isolated text. This position paper argues that toxicity detection should be evaluated as the contextual measurement of situated communicative harm, rather than as single-label text classification. Toxicity is not contained in words alone; it emerges when a communicative act is interpreted by an audience within a normative and social context. We introduce the Contextual Stress Framework (CSF), which defines toxicity as a relation between perceived norm violation and induced stress or disruption. CSF explains why text-intrinsic detectors overflag dialectal or reclaimed language, miss coded or pragmatic abuse, and remain brittle under meaning-preserving transformations. We propose CSF-Eval, an evaluation agenda that separates text risk, norm violation, disruption, uncertainty, and policy action.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。