用多智能体模型自训练,低资源下高效识别网络辱骂内容
Multi-Agent VLMs Guided Self-Training with PNU Loss for Low-Resource Offensive Content Detection
- 多智能体视觉语言模型协同生成伪标签,提升标注可靠性
- 引入PNU损失函数,在少量标注数据下性能接近大模型
- 适合资源有限但需精准识别辱骂内容的平台安全场景
社交媒体中准确检测辱骂内容需要高质量标注数据,但因辱骂样本稀少且人工标注成本高,常面临数据稀缺问题。为此,我们提出一种自训练框架,利用大量未标注数据通过多智能体视觉语言模型(MA-VLMs)协作生成伪标签。初始轻量级分类器基于少量标注数据训练,随后在未标注数据上迭代生成伪标签:分类器与MA-VLMs意见一致的样本构成‘共识未知’集,分歧样本则为‘冲突未知’集。为增强标签可信度,MA-VLMs模拟监管者与用户双视角,捕捉规范性与主观性双重立场。分类器采用新型正-负-未标记(PNU)损失函数,联合优化已标注、共识未知与冲突未知数据,有效缓解伪标签噪声。在基准数据集上的实验表明,该框架在有限监督下显著优于基线方法,并逼近大规模模型性能。
原文摘要 · Abstract (English)
Accurate detection of offensive content on social media demands high-quality labeled data; however, such data is often scarce due to the low prevalence of offensive instances and the high cost of manual annotation. To address this low-resource challenge, we propose a self-training framework that leverages abundant unlabeled data through collaborative pseudo-labeling. Starting with a lightweight classifier trained on limited labeled data, our method iteratively assigns pseudo-labels to unlabeled instances with the support of Multi-Agent Vision-Language Models (MA-VLMs). Un-labeled data on which the classifier and MA-VLMs agree are designated as the Agreed-Unknown set, while conflicting samples form the Disagreed-Unknown set. To enhance label reliability, MA-VLMs simulate dual perspectives, moderator and user, capturing both regulatory and subjective viewpoints. The classifier is optimized using a novel Positive-Negative-Unlabeled (PNU) loss, which jointly exploits labeled, Agreed-Unknown, and Disagreed-Unknown data while mitigating pseudo-label noise. Experiments on benchmark datasets demonstrate that our framework substantially outperforms baselines under limited supervision and approaches the performance of large-scale models
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。