探索跨语言文本净化中关键词引导的实效与代价
AEGIS: Awareness-Enhanced Guidance for Iterative Safeguard

- 用关键词级提示+冻结生成模型,结构化引导改写
- 发现关键词引导能改变毒性与语义保留的权衡关系
- 适合研究多语言净化中可控性机制的研究者
段落级理由常被视为提升文本净化可控制性的手段,但其在何时有效、何时带来权衡仍不明确。本文提出意识增强型迭代防护框架AEGIS,用于研究英文、中文和韩文的段落级引导多语言净化。AEGIS结合段落级检测器输出与冻结的生成模型主干,允许在重写过程中提供有害片段、强度标签及目标属性作为结构化引导。我们并未宣称达到最先进的净化性能,而是分析段落引导如何影响不同生成模型家族、模型规模和语言下的毒性降低与语义保留之间的平衡。结果表明,段落引导的净化效果具有条件性:显式理由会改变毒性降低与语义保留的权衡,但其影响强烈依赖于生成模型主干和语言上下文。这些发现揭示了段落级控制信号在多语言净化中的潜力与局限。
原文摘要 · Abstract (English)
Span-level rationales are often assumed to improve controllability in text detoxification, but it remains unclear when such guidance helps and when it introduces trade-offs. We present Awareness-Enhanced Guidance for Iterative Safeguard (AEGIS) as an exploratory framework for studying span-guided multilingual detoxification across English, Mandarin Chinese, and Korean. AEGIS combines span-level detector outputs with frozen generator backbones, allowing harmful spans, intensity labels, and target attributes to be provided as structured guidance during rewriting. Rather than claiming state-of-the-art detoxification performance, we analyze how span guidance affects the balance between toxicity reduction and meaning preservation across generator families, model scales, and languages. Our results suggest that span-guided detoxification is conditionally useful: explicit rationales change the trade-off between toxicity reduction and meaning preservation, but their effects depend strongly on the generator backbone and the linguistic context. These findings highlight both the promise and the limitations of span-level control signals for multilingual detoxification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。