arXiv:2512.18027cs.CLcs.CY2025-12被引 2

小模型CoPE实现可调控的内容标注,精度超大模型但仅占1%规模。

CoPE: A Small Language Model for Steerable and Scalable Content Labeling

  • 用矛盾样例训练让模型理解政策而非死记硬背。
  • 在7类有害内容上准确率媲美甚至超越顶尖大模型。
  • 适合需要快速部署、自主可控内容治理的平台使用。

本文介绍了CoPE的方法论,这是一种可政策调控的小型语言模型,能够快速准确地进行内容标注。我们提出一种新型训练方案——矛盾样例训练,使模型学会理解政策而非简单记忆;同时提出二视标注方法,快速构建无歧义的训练数据集。在七个不同危害领域评估中,CoPE的准确率与前沿模型相当或更优,且仅为其1%的参数量。我们公开发布一个90亿参数版本,可在单块消费级显卡上运行。CoPE这类模型标志着分类系统的新范式:将机器学习任务转化为政策撰写任务,为在线平台治理开辟新设计空间。

原文摘要 · Abstract (English)

This paper details the methodology behind CoPE, a policy-steerable small language model capable of fast and accurate content labeling. We present a novel training curricula called Contradictory Example Training that enables the model to learn policy interpretation rather than mere policy memorization. We also present a novel method for generating content policies, called Binocular Labeling, which enables rapid construction of unambiguous training datasets. When evaluated across seven different harm areas, CoPE exhibits equal or superior accuracy to frontier models at only 1% of their size. We openly release a 9 billion parameter version of the model that can be run on a single consumer-grade GPU. Models like CoPE represent a paradigm shift for classifier systems. By turning an ML task into a policy writing task, CoPE opens up new design possibilities for the governance of online platforms.

小模型内容标注政策可控可部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。