构建兼顾社会文化差异的LLM内容审核评估框架,提升评测真实性。
Socio-Culturally Aware Evaluation Framework for LLM-Based Content Moderation
- 基于角色生成多样化内容数据集,模拟真实社交语境。
- 小模型在跨文化内容审核中表现显著下降,暴露其局限性。
- 适合研究公平性、偏见与内容审核的学者与工程师参考。
随着社交媒体和大语言模型的发展,内容审核变得愈发重要。现有许多数据集未能充分代表不同群体,导致评估结果不可靠。为此,我们提出一种面向社会文化背景的内容审核评估框架,并设计了一种基于角色生成的可扩展数据集构建方法。分析显示,该方法生成的数据集能提供更广泛的视角,对LLM构成更大挑战,尤其在小型模型中更为明显,凸显其在处理多样内容时的困难。
原文摘要 · Abstract (English)
With the growth of social media and large language models, content moderation has become crucial. Many existing datasets lack adequate representation of different groups, resulting in unreliable assessments. To tackle this, we propose a socio-culturally aware evaluation framework for LLM-driven content moderation and introduce a scalable method for creating diverse datasets using persona-based generation. Our analysis reveals that these datasets provide broader perspectives and pose greater challenges for LLMs than diversity-focused generation methods without personas. This challenge is especially pronounced in smaller LLMs, emphasizing the difficulties they encounter in moderating such diverse content.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。