构建首个长上下文大模型安全对齐数据集,提升模型在长文本中的安全性。
LongSafety: Enhance Safety for Long-Context LLMs
- 构建包含17,000样本的长上下文安全对齐数据集,平均长度达40.9k token。
- 使用该数据集训练后,模型在长、短上下文中的安全表现均提升,且不损失通用能力。
- 验证了长上下文安全不能简单依赖短上下文数据,新数据集具备跨长度泛化能力。
近期模型架构与长度外推技术的进步显著扩展了大语言模型(LLMs)的上下文长度,推动其在复杂任务中的应用。然而,尽管长上下文LLM能力增强,其安全问题仍缺乏系统研究。尽管短上下文安全对齐已有广泛探索,但长上下文场景下的安全挑战尚未充分解决。本文提出 extbf{LongSafety},一个面向长上下文LLMs的综合性安全对齐数据集,包含10个任务和17,000个样本,平均长度为40.9k tokens。实验表明,使用LongSafety进行训练可同时提升长、短上下文的安全性能,且保持模型的通用能力。进一步证明,长上下文安全并非等同于用短上下文数据对齐的结果,LongSafety具备在不同上下文长度和安全场景下的泛化能力。
原文摘要 · Abstract (English)
Recent advancements in model architectures and length extrapolation techniques have significantly extended the context length of large language models (LLMs), paving the way for their application in increasingly complex tasks. However, despite the growing capabilities of long-context LLMs, the safety issues in long-context scenarios remain underexplored. While safety alignment in short context has been widely studied, the safety concerns of long-context LLMs have not been adequately addressed. In this work, we introduce \textbf{LongSafety}, a comprehensive safety alignment dataset for long-context LLMs, containing 10 tasks and 17k samples, with an average length of 40.9k tokens. Our experiments demonstrate that training with LongSafety can enhance long-context safety performance while enhancing short-context safety and preserving general capabilities. Furthermore, we demonstrate that long-context safety does not equal long-context alignment with short-context safety data and LongSafety has generalizing capabilities in context length and long-context safety scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。