小模型通过问答式设计,实现多模态内容安全分类新突破。
Shieldstral

- 将内容审核转化为二元问答任务,统一多类安全标准
- 30亿参数模型在文本安全上超越近7倍大的模型
- 适合需要轻量化、可定制化安全过滤的部署场景
我们提出Shieldstral,一个30亿参数的策略自适应多模态安全分类器,在文本安全基准上表现与规模接近其7倍的模型相当甚至更优,并在多模态安全分类任务上达到新SOTA。Shieldstral将内容审核建模为二元问答任务,这种简洁范式将多样化的审核任务统一为单一的‘是/否’问题,使具有不同分类体系的异构安全数据集能在同一训练框架下整合。本文提出完整的数据构建方案,涵盖约5410万样本的筛选与生成,以及细粒度评估集以检验策略适应能力。这些设计使小型自适应模型能够媲美甚至超越大型模型。
原文摘要 · Abstract (English)
We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7$\times$ its size on text safety benchmarks and sets a new state of the art on multimodal safety classification. Shieldstral formulates content moderation as a binary question-answering task. This simple formulation unifies diverse moderation tasks into a single yes/no problem, enabling heterogeneous safety datasets with divergent taxonomies to be consolidated under one training framework. We present the data construction recipe, covering curation and generation of approximately 54.1M samples and a fine-grained evaluation set to evaluate policy adaptability. Together, these enable a small adaptive model to match or outperform much larger models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。