arXiv:2502.05291cs.CL2025-02

小模型生成有害内容能力不同,大模型评分不准。

Can LLMs Rank the Harmfulness of Smaller LLMs? We are Not There Yet

  • 用提示词诱发小模型生成各类有害内容
  • 大模型对有害性判断与人类共识仅中低度一致
  • 适合关注AI安全与评估的开发者参考

大型语言模型(LLMs)已广泛应用,理解其风险至关重要。小型模型在计算资源受限的边缘设备上部署时,可能产生不同倾向的有害输出。通常需人工标注有害性以缓解风险,但成本高昂。本文研究两个问题:小型模型在生成有害内容方面如何排序?大型模型能否准确标注?我们通过提示三个小型模型生成歧视性语言、攻击性内容、隐私侵犯或负面引导等有害内容,并收集人类对输出的排序。随后评估三个顶尖大型模型在标注这些输出有害性方面的表现。结果发现,小型模型在有害性上存在差异;大型模型与人类判断的吻合度仅为低至中等。这表明当前仍需进一步研究模型危害缓解方法。

原文摘要 · Abstract (English)

Large language models (LLMs) have become ubiquitous, thus it is important to understand their risks and limitations. Smaller LLMs can be deployed where compute resources are constrained, such as edge devices, but with different propensity to generate harmful output. Mitigation of LLM harm typically depends on annotating the harmfulness of LLM output, which is expensive to collect from humans. This work studies two questions: How do smaller LLMs rank regarding generation of harmful content? How well can larger LLMs annotate harmfulness? We prompt three small LLMs to elicit harmful content of various types, such as discriminatory language, offensive content, privacy invasion, or negative influence, and collect human rankings of their outputs. Then, we evaluate three state-of-the-art large LLMs on their ability to annotate the harmfulness of these responses. We find that the smaller models differ with respect to harmfulness. We also find that large LLMs show low to moderate agreement with humans. These findings underline the need for further work on harm mitigation in LLMs.

大模型安全有害内容评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。