针对中文平台间仇恨评论检测性能下降问题,提出双阈值难例挖掘方法。
Cross-Platform Chinese Offensive Comment Detection via Dual-Threshold Hard Example Mining

- 通过置信度筛选高低置信度难例,构建小规模人工标注集
- 在微博、小红书等四平台测试,性能显著提升
- 适合需要低成本跨平台部署的舆情系统开发者
中文社交媒体的跨平台仇恨评论检测面临性能下降问题。本文首先在COLD数据集上微调干净中文版RoBERTa,建立公平对比的二分类基线。其次构建涵盖微博、小红书、贴吧和知乎的三分类细粒度测试集,使用杰卡德距离和代理距离量化源域与目标域间的领域差异,并系统揭示了基线模型在领域迁移下的性能瓶颈。为此,提出双阈值难例挖掘策略:根据预测置信度从无标签语料中筛选高、低置信度易错样本,仅需少量人工标注的难例,在隐式上下文中进行二次微调,实现低成本跨平台领域自适应。实验表明,优化后的模型在四个平台上均取得显著性能提升。
原文摘要 · Abstract (English)
Cross-platform deployment of offensive comment detection for Chinese social media suffers performance degradation. The paper proposes a dual-threshold hard mining method to address this. First, the clean-Chinese-base RoBERTa is finetuned on COLD to establish a binary baseline for fair comparison. Second, a three-class fine-labeled test set covering Weibo, Xiaohongshu, Tieba, and Zhihu is constructed, domain distances from the source are quantified using Jaccard and Proxy-A Distance, as well as the degradation bottleneck of the baseline under domain shift is systematically revealed. Herein, a dual threshold hard example mining strategy is proposed. High- and low-confidence error-prone samples are filtered from unlabeled corpora by prediction confidence. The model is secondarily finetuned under implicit contexts with merely a small set of manually labeled hard examples, realizing low-cost cross-platform domain adaptation. Experiments reveal significant performance gains of the optimized model across four platforms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。