arXiv:2411.05214cs.CL2024-11被引 11

小模型通过指令微调实现多场景内容审核,效果媲美大模型。

STAND-Guard: A Small Task-Adaptive Content Moderation Model

  • 用指令微调让小模型适配多种内容审核任务
  • 在40多个公开数据集上表现接近GPT-3.5-Turbo
  • 对未见任务的二分类效果接近GPT-4-Turbo

内容审核是保障在线平台安全与负责任大模型发展的关键环节。由于不同审核任务具有特定需求,亟需一种无需大量调优即可准确适应新任务的小模型。本文提出STAND-Guard,一种小型任务自适应内容审核模型。核心思想是:通过对多种内容审核任务进行指令微调,可充分释放小语言模型(SLMs)在未见过(分布外)任务上的潜力。我们还系统研究了训练任务和模型规模对跨任务微调效果的影响。实验表明,STAND-Guard在超过40个公开数据集及真实业务场景的专有数据集上表现与GPT-3.5-Turbo相当;尤为突出的是,在未见的英文二分类任务上,其性能几乎等同于GPT-4-Turbo。

原文摘要 · Abstract (English)

Content moderation, the process of reviewing and monitoring the safety of generated content, is important for development of welcoming online platforms and responsible large language models. Content moderation contains various tasks, each with its unique requirements tailored to specific scenarios. Therefore, it is crucial to develop a model that can be easily adapted to novel or customized content moderation tasks accurately without extensive model tuning. This paper presents STAND-GUARD, a Small Task-Adaptive coNtent moDeration model. The basic motivation is: by performing instruct tuning on various content moderation tasks, we can unleash the power of small language models (SLMs) on unseen (out-of-distribution) content moderation tasks. We also carefully study the effects of training tasks and model size on the efficacy of cross-task fine-tuning mechanism. Experiments demonstrate STAND-Guard is comparable to GPT-3.5-Turbo across over 40 public datasets, as well as proprietary datasets derived from real-world business scenarios. Remarkably, STAND-Guard achieved nearly equivalent results to GPT-4-Turbo on unseen English binary classification tasks

内容审核小模型指令微调任务自适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。