arXiv:2606.27629cs.CLcs.AI2026-06

针对中文平台间仇恨评论检测性能下降问题,提出双阈值难例挖掘方法。

Cross-Platform Chinese Offensive Comment Detection via Dual-Threshold Hard Example Mining

论文配图:Cross-Platform Chinese Offensive Comment Detection via Dual-Threshold Hard Example Mining
图 1 · 摘自论文原文
  • 通过置信度筛选高低置信度难例,构建小规模人工标注集
  • 在微博、小红书等四平台测试,性能显著提升
  • 适合需要低成本跨平台部署的舆情系统开发者

中文社交媒体的跨平台仇恨评论检测面临性能下降问题。本文首先在COLD数据集上微调干净中文版RoBERTa,建立公平对比的二分类基线。其次构建涵盖微博、小红书、贴吧和知乎的三分类细粒度测试集,使用杰卡德距离和代理距离量化源域与目标域间的领域差异,并系统揭示了基线模型在领域迁移下的性能瓶颈。为此,提出双阈值难例挖掘策略:根据预测置信度从无标签语料中筛选高、低置信度易错样本,仅需少量人工标注的难例,在隐式上下文中进行二次微调,实现低成本跨平台领域自适应。实验表明,优化后的模型在四个平台上均取得显著性能提升。

原文摘要 · Abstract (English)

Cross-platform deployment of offensive comment detection for Chinese social media suffers performance degradation. The paper proposes a dual-threshold hard mining method to address this. First, the clean-Chinese-base RoBERTa is finetuned on COLD to establish a binary baseline for fair comparison. Second, a three-class fine-labeled test set covering Weibo, Xiaohongshu, Tieba, and Zhihu is constructed, domain distances from the source are quantified using Jaccard and Proxy-A Distance, as well as the degradation bottleneck of the baseline under domain shift is systematically revealed. Herein, a dual threshold hard example mining strategy is proposed. High- and low-confidence error-prone samples are filtered from unlabeled corpora by prediction confidence. The model is secondarily finetuned under implicit contexts with merely a small set of manually labeled hard examples, realizing low-cost cross-platform domain adaptation. Experiments reveal significant performance gains of the optimized model across four platforms.

仇恨评论检测跨平台迁移难例挖掘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。