通过社区化多智能体框架提升隐性仇恨言论检测效果
Improving Implicit Hate Speech Detection via a Community-Driven Multi-Agent Framework
- 构建包含中心管理员与动态社区代理的多智能体系统
- 在ToxiGen数据集上超越现有提示方法,准确率与公平性双提升
- 引入平衡准确率评估公平性,适合平台内容安全场景
本文提出一种情境化的隐性仇恨言论检测框架,采用多智能体系统,包含一个中心管理员智能体和代表特定群体的动态社区智能体。该方法显式整合公开知识源中的社会文化背景信息,实现身份敏感的内容审核,在具有挑战性的ToxiGen数据集上优于零样本提示、少样本提示及思维链提示等现有方法。通过引入平衡准确率作为分类公平性的核心指标,综合考量真正例率与真负例率的权衡,验证了该社区驱动的协作框架在所有目标群体中均显著提升了分类准确率与公平性。
原文摘要 · Abstract (English)
This work proposes a contextualised detection framework for implicitly hateful speech, implemented as a multi-agent system comprising a central Moderator Agent and dynamically constructed Community Agents representing specific demographic groups. Our approach explicitly integrates socio-cultural context from publicly available knowledge sources, enabling identity-aware moderation that surpasses state-of-the-art prompting methods (zero-shot prompting, few-shot prompting, chain-of-thought prompting) and alternative approaches on a challenging ToxiGen dataset. We enhance the technical rigour of performance evaluation by incorporating balanced accuracy as a central metric of classification fairness that accounts for the trade-off between true positive and true negative rates. We demonstrate that our community-driven consultative framework significantly improves both classification accuracy and fairness across all target groups.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。