arXiv:2607.23175cs.CLcs.AI2026-07

不重训练就能按用户偏好调节模型毒性敏感度。

Beyond a Global Norm: Personalizing Toxicity Sensitivity in Language Models Without Retraining

  • 在推理阶段通过提示、解码、重排序三阶段干预实现个性化
  • 相较基准降低28%-47%的毒性感知偏差,效果显著
  • 适合需要灵活适配不同用户安全偏好的场景

减少毒性常被视为全局对齐问题,但有害语言的判断具有主观性和情境依赖性。本文首次对比评估了三种无需重训练的推理时干预方法:预解码(提示条件化与重写)、解码中(词元、逻辑值和表示引导)以及后解码(候选重排序),用于将语言生成对齐至用户特定的毒性敏感度。基于PRISM数据集构建的毒性敏感度目标进行评估,所有方法均使对齐误差降低28%-47%。然而结果揭示了对齐效果、个性化程度与通用语言质量之间的根本权衡,表明毒性敏感度对齐本质上是多目标问题。

原文摘要 · Abstract (English)

Reducing toxicity is often framed as a global alignment problem, yet perceptions of harmful language are subjective and context-dependent. We present the first comparative evaluation of training-free methods for aligning language generation to user-specific toxicity sensitivities across three inference-time intervention stages: pre-decoding (prompt conditioning and rewriting), in-decoding (token, logit, and representation steering), and post-decoding (candidate re-ranking). Evaluated against toxicity sensitivity targets derived from the PRISM dataset, all methods reduce alignment error by 28-47%. However, the results reveal a fundamental trade-off between alignment effectiveness, personalization, and general language quality, showing how toxicity sensitivity alignment is an inherently multi-objective problem.

语言模型毒性控制个性化推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。