局部修改未必更安全,大范围重写反而更优,关键看有害程度。
When Does Span-Guided Detoxification Help? Human Preferences and Evaluator Diagnostics in a Controlled Comparison

- 按标注片段局部改写,避免无谓改动
- 轻度有害内容用全局重写效果更好
- 需分别评估残留危害与过度修改风险
Span-guided rewriting通过限定修改范围以保留原意,但可能无法充分消除有害意图。我们在混合来源的英文评测集上进行受控对比,包含人工标注输入和HateXplain测试项,在单一生成器固定设置下开展密集盲评。人类偏好显示存在权衡:当局部修改能保留原始立场且不冗余时更受青睐;而更广泛的重写在实现完全缓解时更优。这一差异在不同严重性分层中表现显著——强危害组两种策略竞争激烈,弱危害组则明显偏爱无引导重写。理性注释揭示其根源在于互补的失败风险:局部修改后残留危害,全局修改则导致过度修正。我们视自动评估为诊断工具而非人评替代。毒性-相似性标量、多生成器分析及两个通用LLM裁判再现了整体趋势,但未能复现分层对比。这些情境特异性发现不能确立按严重性路由的规则,反而呼吁建立独立评估缓解充分性和语义保留性的评价协议,并同时报告残留危害与过度修改情况。
原文摘要 · Abstract (English)
Span-guided rewriting aims to preserve meaning by localizing edits to annotated harmful spans, but the same constraint can leave harmful intent insufficiently mitigated. We present a controlled exploratory comparison of span-guided and unguided detoxification on a mixed-source English evaluation set comprising manually curated inputs and HateXplain test items. We conduct a dense blinded human evaluation under a fixed single-generator setting. Human preferences reveal a trade-off rather than a uniformly superior rewriting strategy. Span-guided outputs are favored when localized editing preserves the original stance and avoids unnecessary modification, whereas unguided outputs are favored when broader rewriting achieves more complete mitigation. This contrast varies substantially across the study-defined strata: the two strategies are competitive in the strong stratum, while unguided rewriting is clearly preferred in the mild stratum. Rationale annotations trace this difference to complementary failure risks: residual harm after localized editing and over-modification after broader rewriting. We treat automatic evaluation as a diagnostic rather than a substitute for human judgment. Toxicity-similarity scalarizations, a multi-generator analysis, and two general-purpose LLM judges reproduce parts of the aggregate tendency but do not yield an analogous stratified contrast. These setting-specific findings do not establish a severity-based routing rule. Instead, they motivate evaluation protocols that assess mitigation sufficiency and meaning preservation separately and report both residual harm and over-modification alongside aggregate scores.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。