arXiv:2606.27210cs.CL2026-06被引 1

让大模型理解用户意图,能显著提升安全分类效果。

Paved with True Intents: Intent-Aware Training Improves LLM Safety Classification Across Training Regimes

论文配图:Paved with True Intents: Intent-Aware Training Improves LLM Safety Classification Across Training Regimes
图 1 · 摘自论文原文
  • 将用户意图作为显式信号融入训练,改进安全分类
  • 在5个外部基准上,意图感知模型平均表现最优
  • 适合关注模型安全与推理效率平衡的研究者

我们主张安全分类器应将用户意图作为提示与标签之间的显式信号。为此,我们构建了AIMS数据集,包含1,724个高难度安全提示,每个都配有意图描述和伤害标签。利用AIMS,我们在监督微调、偏好学习、推理蒸馏和强化学习等多种训练范式下评估了意图感知训练的效果。尽管数据集规模有限,但其支持的模型在各类训练中均表现优异:基于模型生成意图错误的DPO优于SFT;意图条件蒸馏在多数师生对中胜过仅依赖推理的蒸馏。尤为突出的是,通过GRPO直接奖励意图忠实性,在五个外部安全基准上实现最强平均性能,且意图感知模型在推理延迟与F1值之间形成帕累托前沿。结果表明,忠实建模意图是提升安全分类鲁棒性的高效监督信号。

原文摘要 · Abstract (English)

We argue that safety classifiers should model user intent as an explicit signal between the prompt and the final label. To study this, we introduce AIMS, a human-annotated dataset of 1,724 difficult safety prompts, each paired with an intent description and harm label. We use AIMS to evaluate intent-aware training across supervised fine-tuning, preference learning, reasoning distillation, and reinforcement learning. Despite its size, AIMS enables competitive safety classifiers across training regimes: DPO from model-generated intent errors improves over SFT, and intent-conditioned distillation outperforms reasoning-only distillation in most teacher-student pairs. Most notably, directly rewarding intent faithfulness with GRPO yields the strongest average performance across five external safety benchmarks, while our intent-aware models form the inference latency-F1 Pareto frontier. These results show that faithful intent modeling is a compact, high-quality supervision signal for more robust safety classifiers.

安全分类意图建模强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。