arXiv:2608.15600cs.AI2026-08

构建可验证的中文辱骂内容审核基准,让决策有据可查。

VARM-Bench: Benchmarking Verifiable Structured Reasoning in Chinese Abusive Speech Moderation

论文配图:VARM-Bench: Benchmarking Verifiable Structured Reasoning in Chinese Abusive Speech Moderation
图 1 · 摘自论文原文
  • 为中文辱骂内容审核设计带明确锚点的链式推理路径。
  • 模型标签准确率高但完整决策记录错误率仍显著。
  • 适合关注内容审核透明性与可审计性的研究者使用。

网络辱骂内容泛滥加剧了中文社交媒体文本审核的可靠性需求。现有中文基准支持标签分类、细粒度毒性归类和目标感知提取,但缺乏对审核决策依据的统一可验证表示。本文提出 VARM-Bench,一个面向中文辱骂内容审核的场域锚定链式推理基准。每条样本包含简洁自然语言的推理过程,并对六个决策项(目标、目标类型、目标明确性、作者立场、危害性标签、细粒度类别)提供显式锚点。采用确定性评估协议,不依赖大模型裁判,检验字段正确性、目标对齐性、输出有效性、完整记录一致性及基于正确最终决策的隐藏记录错误。在统一结构化输出协议下,我们通过零样本提示、分类体系引导和结构化思维链监督,评估多个模型家族的语言模型,并分析词汇线索敏感性和字段级错误。结果表明,高标签级性能可能掩盖完整审核记录中的显著错误。VARM-Bench 提供了一个可审计、可复现的基准,用于评估中文辱骂内容审核中可验证的推理依据。

原文摘要 · Abstract (English)

The widespread circulation of abusive online content has increased the need for reliable moderation of Chinese social-media text. Existing Chinese benchmarks support label classification, fine-grained toxicity categorization, and target-aware extraction, but do not provide a unified representation for deterministically verifying the stated basis of a moderation decision. We introduce VARM-Bench, a benchmark for field-anchored chain-of-thought rationales in Chinese abusive-speech moderation. Each instance contains a concise natural-language rationale with explicit anchors for six decisions: target, target type, target explicitness, author stance, harmfulness label, and fine-grained category. Our deterministic protocol evaluates field correctness, target alignment, output validity, complete-record agreement, and hidden record errors conditioned on correct final decisions, without relying on an LLM judge. Under a common structured-output protocol, we evaluate language models across multiple model families using zero-shot prompting, taxonomy guidance, and structured CoT supervision, and analyze lexical-cue sensitivity and field-level errors. Results show that strong label-level performance can conceal substantial errors in complete moderation records. VARM-Bench provides an auditable and reproducible benchmark for evaluating verifiable moderation rationales in Chinese abusive-speech moderation.

内容审核可解释性中文NLP链式推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。