arXiv:2602.02536cs.LGcs.AI2026-02被引 1

用多维度推理轨迹提升多模态内容安全审核效果

From Sparse Decisions to Dense Reasoning: A Multi-attribute Trajectory Paradigm for Multimodal Moderation

  • 构建包含证据锚定到响应生成的完整决策路径
  • 仅用40%数据达领先模型性能,多模态评测新基准
  • 适合需要可解释性与高精度的安全审核场景

安全审核对识别有害内容至关重要。尽管文本安全审核已取得成功,其多模态版本仍受限于数据与监督信号的双重稀疏性。传统二值标签导致模型学习捷径,掩盖了有效区分所需的内在分类边界。为此,我们提出一种新范式(UniMod),将稀疏决策转化为密集推理轨迹。通过构建涵盖证据锚定、模态评估、风险映射、策略决策和响应生成的结构化轨迹,将单一决策任务重构为多维边界学习过程。该方法迫使模型基于显式安全语义做判断,避免收敛至表面捷径。为此,我们开发了多头标量奖励模型(UniRM),在响应生成阶段提供属性级评分作为多维监督。此外,引入专用优化策略解耦任务特定参数并平衡训练动态,有效缓解多任务学习中的目标干扰。实验证明,UniMod在文本审核上表现相当,并以少于40%的训练数据量创下多模态审核新基准。消融实验进一步验证了多属性轨迹推理的有效性,为多模态审核提供了高效且可解释的框架。

原文摘要 · Abstract (English)

Safety moderation is pivotal for identifying harmful content. Despite the success of textual safety moderation, its multimodal counterparts remain hindered by a dual sparsity of data and supervision. Conventional reliance on binary labels lead to shortcut learning, which obscures the intrinsic classification boundaries necessary for effective multimodal discrimination. Hence, we propose a novel learning paradigm (UniMod) that transitions from sparse decision-making to dense reasoning traces. By constructing structured trajectories encompassing evidence grounding, modality assessment, risk mapping, policy decision, and response generation, we reformulate monolithic decision tasks into a multi-dimensional boundary learning process. This approach forces the model to ground its decision in explicit safety semantics, preventing the model from converging on superficial shortcuts. To facilitate this paradigm, we develop a multi-head scalar reward model (UniRM). UniRM provides multi-dimensional supervision by assigning attribute-level scores to the response generation stage. Furthermore, we introduce specialized optimization strategies to decouple task-specific parameters and rebalance training dynamics, effectively resolving interference between diverse objectives in multi-task learning. Empirical results show UniMod achieves competitive textual moderation performance and sets a new multimodal benchmark using less than 40\% of the training data used by leading baselines. Ablations further validate our multi-attribute trajectory reasoning, offering an effective and efficient framework for multimodal moderation. Supplementary materials are available at \href{https://trustworthylab.github.io/UniMod/}{project website}.

内容审核多模态推理轨迹强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。