用大模型将仇恨言论转为无攻击性表达,保留原意。
Detoxify: A framework for abusive text transformation using LLMs
- 用大模型识别并重构仇恨言论与脏话
- GPT-4o与DeepSeek效果相近,Groq表现差异显著
- 适合内容安全、AI伦理方向研究者参考
尽管大型语言模型在自然语言处理任务中取得显著进展,但在识别和转换仇恨言论与粗俗内容方面仍需探索。本文提出Detoxify框架,利用大模型将含仇恨言论和脏话的推文及评论转化为无攻击性文本,同时保持原意。评估了Gemini、GPT-4o、DeepSeek和Groq四种前沿大模型的识别能力,并对原始与转化后数据进行情感与语义分析。结果显示,相较于其他模型,Groq表现迥异,常过度使用正面表述重构句子,导致原上下文丢失或改变;GPT-4o与DeepSeek则表现出相似效果。
原文摘要 · Abstract (English)
Although Large Language Models (LLMs) have demonstrated significant advancements in natural language processing tasks, their effectiveness in the classification and transformation of abusive text into non-abusive versions remains an area for exploration. In this study, we present Detoxify: a framework that employs LLMs to transform abusive text (tweets and reviews) containing hate speech and profanity into non-abusive text while retaining the original intent. We evaluate the performance of four state-of-the-art LLMs, such as Gemini, GPT-4o, DeekSeek and Groq, on their ability to identify abusive text. We aim to transform and obtain a text that is clean of abusive and inappropriate content, but maintains a similar level of sentiment and semantics, i.e. the transformed text needs to maintain its message. Afterwards, we evaluate the raw and transformed datasets with sentiment analysis and semantic analysis. Our results show Groq provides vastly different results when compared with other LLMs. We have identified similarities between GPT-4o and DeepSeek. Groq stood out as the most distinct, as it often restructured sentences with excessive positive phrasing, with the original context lost or altered.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。