利用标注分歧信号提升毒性内容检测的准确性和可解释性。
A Collaborative Content Moderation Framework for Toxicity Detection based on Conformalized Estimates of Annotation Disagreement
- 多任务学习联合预测毒性与标注分歧,捕捉内容模糊性。
- 结合置信区间估计,提升模型对不确定性的识别能力。
- 支持人工审核阈值调整,适用于高敏感内容治理场景。
内容审核通常结合人工与机器学习模型。然而,这类系统常依赖存在显著标注分歧的数据,反映出毒性判断的主观性。我们不将分歧视为噪声,而是将其视为内容固有模糊性的有效信号。本文提出一种新型内容审核框架,强调捕捉标注分歧的重要性。方法采用多任务学习:以毒性分类为主任务,标注分歧为辅助任务;同时引入置信区间估计(Conformal Prediction)技术,量化评论标注的模糊性及模型对毒性与分歧预测的不确定性。框架还允许审核员调节分歧阈值,灵活决定何时触发人工复审。实验表明,该联合方法在性能、校准度和不确定性估计方面均优于单任务模型,且参数效率更高,优化了审核流程。
原文摘要 · Abstract (English)
Content moderation typically combines the efforts of human moderators and machine learning models. However, these systems often rely on data where significant disagreement occurs during moderation, reflecting the subjective nature of toxicity perception. Rather than dismissing this disagreement as noise, we interpret it as a valuable signal that highlights the inherent ambiguity of the content,an insight missed when only the majority label is considered. In this work, we introduce a novel content moderation framework that emphasizes the importance of capturing annotation disagreement. Our approach uses multitask learning, where toxicity classification serves as the primary task and annotation disagreement is addressed as an auxiliary task. Additionally, we leverage uncertainty estimation techniques, specifically Conformal Prediction, to account for both the ambiguity in comment annotations and the model's inherent uncertainty in predicting toxicity and disagreement.The framework also allows moderators to adjust thresholds for annotation disagreement, offering flexibility in determining when ambiguity should trigger a review. We demonstrate that our joint approach enhances model performance, calibration, and uncertainty estimation, while offering greater parameter efficiency and improving the review process in comparison to single-task methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。