提出新指标RDC,对比多款大模型的伦理与安全差距。
A Comparative Analysis of Ethical and Safety Gaps in LLMs using Relative Danger Coefficient
- 引入相对危险系数RDC量化大模型潜在危害
- 对比DeepSeek-V3、GPT、Gemini等模型的安全表现
- 强调高风险场景需强化人工监督
近年来,人工智能与大语言模型在自然语言理解与生成方面取得显著进展,但其快速发展也引发关于安全、滥用、歧视及社会影响的伦理关切。本文对多种AI模型——包括最新版DeepSeek-V3(带推理与不带推理版本)、GPT系列(4o、3.5 Turbo、4 Turbo、o1/o3 mini)以及Gemini(1.5 flash、2.0 flash和2.0 flash exp)——进行了伦理性能的比较分析,强调在高风险情境下需建立强有力的监管机制。此外,本文提出一种新的模型危害评估指标——相对危险系数(Relative Danger Coefficient, RDC),用于系统衡量大语言模型可能造成的实际危害。
原文摘要 · Abstract (English)
Artificial Intelligence (AI) and Large Language Models (LLMs) have rapidly evolved in recent years, showcasing remarkable capabilities in natural language understanding and generation. However, these advancements also raise critical ethical questions regarding safety, potential misuse, discrimination and overall societal impact. This article provides a comparative analysis of the ethical performance of various AI models, including the brand new DeepSeek-V3(R1 with reasoning and without), various GPT variants (4o, 3.5 Turbo, 4 Turbo, o1/o3 mini) and Gemini (1.5 flash, 2.0 flash and 2.0 flash exp) and highlights the need for robust human oversight, especially in situations with high stakes. Furthermore, we present a new metric for calculating harm in LLMs called Relative Danger Coefficient (RDC).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。