模型越追求搞笑,越容易生成刻板与有毒内容。
Engagement Undermines Safety: How Stereotypes and Toxicity Shape Humor in Language Models
- 通过幽默评分联合检测模型输出的刻板与毒性。
- 有害内容幽默分高出10%-21%,角色提示下更严重。
- 适合关注AI安全与内容风险的研究者阅读。
大型语言模型在创意写作和互动内容中的应用日益广泛,引发安全担忧。本文以幽默生成为测试场景,评估现代LLM流程中趣味性优化如何与有害内容耦合,同时测量幽默度、刻板印象和毒性。结合信息论指标分析不协调信号,发现六种模型中,有害输出获得更高幽默评分,且在角色提示下进一步上升,表明生成器与评价者间存在偏见放大循环。信息论分析显示,有害线索扩大预测不确定性,甚至使某些模型认为有害笑点更可预期,暗示有害内容嵌入幽默分布结构中。外部验证在讽刺生成任务中显示,模型生成的讽刺内容增加刻板与毒性,包括闭源模型。定量结果表明:刻板/有毒笑话平均幽默分提升10%-21%;在模型标记有趣的笑话中,刻板笑话出现频率高11%-28%;人类判断中也多出最多10%。
原文摘要 · Abstract (English)
Large language models are increasingly used for creative writing and engagement content, raising safety concerns about the outputs. Therefore, casting humor generation as a testbed, this work evaluates how funniness optimization in modern LLM pipelines couples with harmful content by jointly measuring humor, stereotypicality, and toxicity. This is further supplemented by analyzing incongruity signals through information-theoretic metrics. Across six models, we observe that harmful outputs receive higher humor scores which further increase under role-based prompting, indicating a bias amplification loop between generators and evaluators. Information-theoretic analyses show harmful cues widen predictive uncertainty and surprisingly, can even make harmful punchlines more expected for some models, suggesting structural embedding in learned humor distributions. External validation on an additional satire-generation task with human perceived funniness judgments shows that LLM satire increases stereotypicality and typically toxicity, including for closed models. Quantitatively, stereotypical/toxic jokes gain $10-21\%$ in mean humor score, stereotypical jokes appear $11\%$ to $28\%$ more often among the jokes marked funny by LLM-based metric and up to $10\%$ more often in generations perceived as funny by humans.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。