arXiv:2501.13952cs.CLcs.AI2025-01被引 3

解决大模型安全与实用的矛盾,提升化学问答表现

The Dual-use Dilemma in LLMs: Do Empowering Ethical Capacities Make a Degraded Utility?

  • 用DPO框架优化伦理与实用平衡,生成31.6k化学问答三元组数据
  • 在安全与实用双指标上超越GPT-4o、Claude-3等主流模型
  • 揭示长思维链模型存在伦理漏洞,适合模型安全研究者参考

近年来,大型语言模型(LLMs)在多个领域持续优化,但其伦理风险与实用价值之间的权衡问题仍被忽视。本文提出一种基于直接偏好优化(DPO)的对齐框架,以化学领域应用为验证场景,系统缓解安全与效用的冲突。通过GPT辅助的三阶段数据生成流程,构建包含31.6万条三元组实例的LibraChemQA数据集,并引入平衡种子机制,兼顾合法与非法请求。框架还设计重述机制实现高效数据增强,提升模型化学理解能力。我们开发了一种结合大模型评判的混合评估方案,精确衡量安全与效用。实验表明,该模型在综合性能上显著优于Claude-3、GPT-4o和LLaMA-3,分别领先13.44%、7.16%和7.10%。此外,对DeepSeek-R1的测试揭示其长链思维(CoT)推理过程带来严重伦理风险,提示衍生模型同样存在隐患。

原文摘要 · Abstract (English)

Recent years have witnessed extensive efforts to enhance Large Language Models (LLMs) across various domains, alongside growing attention to their ethical implications. However, a critical challenge remains largely overlooked: LLMs must balance between rejecting harmful requests for safety and accommodating legitimate ones for utility. This paper presents a Direct Preference Optimization (DPO) based alignment framework that achieves better overall performance by addressing this ethical-utility trade-off, using chemical domain applications as a proof-of-concept. Our alignment pipeline starts with a GPT-assisted three-phase data generation scheme, in which we create LibraChemQA, a chemical question-answering dataset comprising 31.6k triplet instances. By incorporating an innovative balanced seed in the data generation process, our framework systematically considers both legitimate and illegitimate requests. The framework also introduces a rephrasing mechanism for efficient data augmentation that enhances the model's chemical comprehension. We further develop a novel hybrid evaluation scheme with LLM judges for precise assessment of both safety and utility. Experimental results demonstrate our model's substantial improvements in overall performance where both safety and utility are considered - the resulting model outperforms leading LLMs including Claude-3, GPT-4o, and LLaMA-3 by margins of 13.44%, 7.16%, and 7.10% respectively on our released benchmark. At the end of this paper, we analyze experimental results obtained from testing DeepSeek-R1 on our benchmark and reveal the critical ethical concerns raised by this highly acclaimed model. We highlight that the long Chain-of-Thought (CoT) reasoning process employed by DeepSeek-R1, as well as other LLMs distilled from it, introduces significant ethical vulnerabilities when exposed to users.

大模型对齐伦理安全化学问答数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。