arXiv:2504.01903cs.CLcs.AI2025-04AAAI被引 64

用1000条数据提升大模型推理安全,副作用极小。

STAR-1: Safer Alignment of Reasoning LLMs with 1K Data

  • 从多元数据出发,结合政策引导生成深度思考样本。
  • 微调后安全性能平均提升40%,推理能力仅下降1.1%。
  • 适合关注大模型安全对齐的研究者与工程师。

本文提出STAR-1,一个专为大型推理模型(LRMs)如DeepSeek-R1设计的高质量、仅1000条规模的安全数据集。基于多样性、深度推理与严格筛选三大原则,首先整合多个开源安全数据源,再通过安全策略生成有依据的推理样本,最后利用GPT-4o构建的安全评分系统筛选符合最佳实践的训练样本。实验表明,使用STAR-1微调LRMs,在四个基准上平均安全性能提升40%,而五项推理任务的平均能力仅下降1.1%。消融实验证实了设计原则的有效性,并验证其在LRMs及传统LLMs中的适用性。项目主页:https://ucsc-vlaa.github.io/STAR-1。

原文摘要 · Abstract (English)

This paper introduces STAR-1, a high-quality, just-1k-scale safety dataset specifically designed for large reasoning models (LRMs) like DeepSeek-R1. Built on three core principles -- diversity, deliberative reasoning, and rigorous filtering -- STAR-1 aims to address the critical needs for safety alignment in LRMs. Specifically, we begin by integrating existing open-source safety datasets from diverse sources. Then, we curate safety policies to generate policy-grounded deliberative reasoning samples. Lastly, we apply a GPT-4o-based safety scoring system to select training examples aligned with best practices. Experimental results show that fine-tuning LRMs with STAR-1 leads to an average 40% improvement in safety performance across four benchmarks, while only incurring a marginal decrease (e.g., an average of 1.1%) in reasoning ability measured across five reasoning tasks. Extensive ablation studies further validate the importance of our design principles in constructing STAR-1 and analyze its efficacy across both LRMs and traditional LLMs. Our project page is https://ucsc-vlaa.github.io/STAR-1.

模型安全推理对齐数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。