测试视觉语言模型如何根据规则链推理内容审核决策。
RuleSafe-VL: Evaluating Rule-Conditioned Decision Reasoning in Vision-Language Content Moderation

- 构建93条原子规则与92种规则关系,模拟真实平台审核逻辑。
- 10个主流模型在规则关联恢复上最高仅64.8分,安全模型最低不足7分。
- 适合研究模型可解释性与可信AI的学者使用。
平台内容审核依赖明确的政策规则和上下文条件来决定内容是否允许、限制或删除。正确判断需依赖激活的规则、规则间的交互关系以及证据充分性。当前多模态安全基准大多简化为匹配最终标签,未检验背后的规则结构。为此,我们提出RuleSafe-VL,一个用于视觉-语言内容审核中规则条件决策推理的基准。基于公开平台审核政策,其形式化了93条原子规则与92种规则关系,生成2,166个高风险政策族下的情境敏感图文案例。四个诊断任务将审核过程分解为规则激活识别、规则交互恢复、决策充分性判断及缺失上下文补全后的结果解析。对10个前沿开源与安全导向视觉语言模型的实验显示,规则关系恢复是主要瓶颈,最佳模型宏平均F1仅为64.8,部分安全模型低于7。决策状态预测也不可靠,最高达64.5。RuleSafe-VL推动审核评估从最终标签评分转向规则条件决策推理的诊断分析。
原文摘要 · Abstract (English)
Platform content moderation applies explicit policy rules and context-dependent conditions to decide whether user content is allowed, restricted, or removed. A correct moderation outcome must therefore depend on which rules a case activates, how those rules interact, and whether the available evidence is sufficient. Current multimodal safety benchmarks largely reduce moderation to matching predefined final labels, leaving this underlying rule structure untested. As a result, a high benchmark score reveals little about whether a model applies the policy correctly or arrives at the correct label through superficial cues. To evaluate this rule-governed process, we introduce RuleSafe-VL, a benchmark for rule-conditioned decision reasoning in vision-language content moderation. Derived from publicly available platform moderation policies, RuleSafe-VL formalizes 93 atomic rules and 92 typed rule relations, yielding 2,166 context-sensitive image-text cases across three high-risk policy families. Its four diagnostic tasks decompose moderation into a rule-conditioned decision chain. They identify activated rules, recover rule interactions, judge decision sufficiency, and resolve outcomes once missing context is supplied. Experiments on 10 frontier, open-source, and safety-oriented VLMs reveal rule-relation recovery as the dominant bottleneck, where the best model reaches only 64.8 Macro-F1 and some safety-oriented models fall below 7 Macro-F1. Decision-state prediction also remains unreliable, peaking at 64.5 Macro-F1. RuleSafe-VL shifts moderation evaluation from final-label scoring toward diagnostic assessment of rule-conditioned decision reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。