arXiv:2606.13397cs.HCcs.AI2026-06

用少数群体经验增强AI内容审核,让系统更懂边缘群体的敏感表达。

Mod-Guide: An LLM-based Content Moderation Feedback System to Address Insensitive Speech toward Indigenous Ethnic and Religious Minority Communities

论文配图:Mod-Guide: An LLM-based Content Moderation Feedback System to Address Insensitive Speech toward Indigenous Ethnic and Religious Minority Communities
图 1 · 摘自论文原文
  • 通过社区共建语料库,将少数族裔真实经历融入审核流程。
  • 引入RAG技术后,模型对少数群体话语的回应准确率显著提升。
  • 适合关注AI伦理、文化包容与数字正义的研究者和实践者。

语言既是边缘化的工具,也是抵抗的手段,尤其对在线环境中遭受不敏感言论冲击的少数群体而言。随着内容审核日益依赖大语言模型(LLM),人们担忧这些系统能否识别出忽视或边缘化历史上被代表不足群体文化与宗教视角的隐性伤害——如隐性抹除、误述或规范性表述,而非明显敌意。本文聚焦孟加拉国最大的宗教少数群体印度教徒与原住民族群查克马人,探究基于LLM的审核系统的认识论局限,并探索纳入少数群体视角的方法。我们与社区成员共同构建了一个以文化为基础的不敏感言论语料库,并利用检索增强生成(RAG)技术将其叙事整合进审核流程。所提出的工具Mod-Guide通过引入来自生活经验的上下文线索,提升了LLM对少数群体观点的敏感度。混合方法评估显示,经RAG增强的审核回应在情境准确性上更优,且不同族裔参与者感知差异显著。本研究推动了人机交互、人工智能伦理与社会计算领域的发展,强调修复性正义与诠释性包容在内容审核系统设计中的重要性。

原文摘要 · Abstract (English)

Language operates as a mechanism of both marginalization and resistance, especially for minority communities navigating insensitive and harmful speech online. As content moderation increasingly depends on large language models (LLMs), concerns arise about whether these systems can recognize culturally insensitive speech-language that disregards or marginalizes the cultural and religious perspectives of historically underrepresented communities, often through implicit erasure, misrepresentation, or normative framing, rather than overt hostility. Focusing on Bangladesh's Hindu and Chakma communities -- the country's largest religious and Indigenous ethnic minorities, respectively -- this paper investigates the epistemic limits of LLM-based moderation systems and explores methods for incorporating minority perspectives. We co-created a culturally grounded corpus of insensitive speech with community members and integrated their narratives into moderation pipelines using retrieval augmented generation (RAG). Our tool, Mod-Guide, improves LLM sensitivity to minority viewpoints by leveraging contextual cues derived from lived experience. Through mixed-method evaluations involving both minority and majority participants, we demonstrate that RAG-enhanced moderation responses are more contextually accurate and perceived differently across ethnic lines. This work advances research in human-computer interaction, AI ethics, and social computing by foregrounding restorative justice and hermeneutical inclusion in the design of content moderation systems.

AI伦理内容审核少数群体RAG

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。