arXiv:2505.21710cs.CL2025-05

测试ChatGPT识别网络不当言论的能力,发现其在过滤不雅内容上表现良好,但误判目标攻击性语言较多。

Assessing and Refining ChatGPT's Performance in Identifying Targeting and Inappropriate Language: A Comparative Study

  • 用多轮迭代优化提示词,提升ChatGPT对不当语言的识别准确率
  • 版本6中检测不当内容准确率显著提升,但针对性攻击语句误判率偏高
  • 适合关注AI内容审核优化的研究者与平台安全团队参考

本研究评估了先进自然语言处理模型ChatGPT在识别社交网络评论中针对性语言和不当语言方面的表现。随着社交媒体用户生成内容规模激增,AI在内容审核中的作用日益重要。通过与众包标注和专家评估对比,分析了ChatGPT在准确性、检测范围和一致性上的表现。结果显示,经过多轮迭代优化后,版本6的ChatGPT在不当内容检测上表现优异,准确率明显提升;但在识别针对性语言时存在波动,误报率高于专家判断。研究证明了类似ChatGPT的AI模型在增强自动化内容审核系统方面的潜力,同时也指出了持续优化和提升上下文理解能力的必要性,以更有效地遏制网络有害行为。

原文摘要 · Abstract (English)

This study evaluates the effectiveness of ChatGPT, an advanced AI model for natural language processing, in identifying targeting and inappropriate language in online comments. With the increasing challenge of moderating vast volumes of user-generated content on social network sites, the role of AI in content moderation has gained prominence. We compared ChatGPT's performance against crowd-sourced annotations and expert evaluations to assess its accuracy, scope of detection, and consistency. Our findings highlight that ChatGPT performs well in detecting inappropriate content, showing notable improvements in accuracy through iterative refinements, particularly in Version 6. However, its performance in targeting language detection showed variability, with higher false positive rates compared to expert judgments. This study contributes to the field by demonstrating the potential of AI models like ChatGPT to enhance automated content moderation systems while also identifying areas for further improvement. The results underscore the importance of continuous model refinement and contextual understanding to better support automated moderation and mitigate harmful online behavior.

AI审核语言识别内容安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。