arXiv:2601.06700cs.CLcs.AI2026-01

分析大模型生成有害内容的规律及影响因素

Characterising Toxicity in Generative Large Language Models

  • 通过分析提示词的词汇与语法特征,研究毒性输出的生成机制
  • 发现特定语言模式显著提升模型产生有害内容的概率
  • 为安全对齐提供可量化的干预依据,适合模型安全研究者

近年来,注意力机制推动了自然语言处理(NLP)的快速发展,基于Transformer的解码器架构在文本处理和生成任务中表现卓越。然而,语言模型仍易生成不当、攻击性或有害内容,统称为“毒性输出”。尽管已有如人类反馈强化学习(RLHF)等对齐方法,但这些防护常被精心设计的提示绕过。本文系统评估大模型在不同提示下生成毒性内容的程度,并探究影响此类输出的语言因素——包括词汇与句法特征。研究揭示特定语言模式会显著提高毒性生成概率,为理解模型安全风险提供了实证基础。

原文摘要 · Abstract (English)

In recent years, the advent of the attention mechanism has significantly advanced the field of natural language processing (NLP), revolutionizing text processing and text generation. This has come about through transformer-based decoder-only architectures, which have become ubiquitous in NLP due to their impressive text processing and generation capabilities. Despite these breakthroughs, language models (LMs) remain susceptible to generating undesired outputs: inappropriate, offensive, or otherwise harmful responses. We will collectively refer to these as ``toxic'' outputs. Although methods like reinforcement learning from human feedback (RLHF) have been developed to align model outputs with human values, these safeguards can often be circumvented through carefully crafted prompts. Therefore, this paper examines the extent to which LLMs generate toxic content when prompted, as well as the linguistic factors -- both lexical and syntactic -- that influence the production of such outputs in generative models.

大模型安全毒性检测语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。