GPT-4o-mini会自动弱化敏感内容,无需明确指令。
What's Taboo for You? - An Empirical Evaluation of LLMs Behavior Toward Sensitive Content
- 通过改写敏感文本,模型自发降低语言攻击性。
- 贬损和禁忌词汇使用量显著下降,体现隐式过滤行为。
- 无需训练即可识别敏感语句,适合安全审查场景。
专有大型语言模型(LLMs)表现出礼貌、正式及隐式内容过滤倾向。以往研究多聚焦于显式训练模型进行内容净化,但对模型是否在无明确指令下隐式净化语言仍缺乏探索。本研究实证分析了 GPT-4o-mini 在重写敏感内容时的隐式调节行为,并评估敏感度变化程度。实验表明,GPT-4o-mini 系统性地将内容转向较低敏感类别,显著减少贬损与禁忌用语。同时,我们评估了 LLMs 的零样本敏感度分类能力,对比其与传统方法的表现。
原文摘要 · Abstract (English)
Proprietary Large Language Models (LLMs) have shown tendencies toward politeness, formality, and implicit content moderation. While previous research has primarily focused on explicitly training models to moderate and detoxify sensitive content, there has been limited exploration of whether LLMs implicitly sanitize language without explicit instructions. This study empirically analyzes the implicit moderation behavior of GPT-4o-mini when paraphrasing sensitive content and evaluates the extent of sensitivity shifts. Our experiments indicate that GPT-4o-mini systematically moderates content toward less sensitive classes, with substantial reductions in derogatory and taboo language. Also, we evaluate the zero-shot capabilities of LLMs in classifying sentence sensitivity, comparing their performances against traditional methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。