用历史判例增强多模态内容过滤,让审核系统更灵活自适应。
Customize Multi-modal RAI Guardrails with Precedent-based predictions
- 基于相似样本的判例推理,替代固定规则进行判断
- 在少样本和全量数据下均超越现有方法,支持新政策快速适配
- 适合需要频繁更新审核标准的平台或个性化内容管理场景
多模态审核机制需根据用户自定义策略过滤图像内容,识别仇恨言论、强化刻板印象、包含色情信息或传播虚假信息等。然而在实际应用中,用户常需高度定制化且多样化的策略,却难以提供充足示例。理想审核系统应可扩展至多种策略,并在无需大量重训的前提下适应用户标准变化。现有微调方法通常依赖预设策略,限制了对新策略的泛化能力,或需大量重训;而无训练方法受限于上下文长度,难以全面整合所有策略。为此,本文提出以“判例”(即与输入相似的历史数据点的推理过程)作为模型判断依据,取代固定策略,显著提升系统的灵活性与适应性。本文引入批判-修正机制收集高质量判例,并提出两种利用判例实现鲁棒预测的策略。实验表明,该方法在少样本和全数据场景下均优于已有方法,且对新型政策具有更强泛化能力。
原文摘要 · Abstract (English)
A multi-modal guardrail must effectively filter image content based on user-defined policies, identifying material that may be hateful, reinforce harmful stereotypes, contain explicit material, or spread misinformation. Deploying such guardrails in real-world applications, however, poses significant challenges. Users often require varied and highly customizable policies and typically cannot provide abundant examples for each custom policy. Consequently, an ideal guardrail should be scalable to the multiple policies and adaptable to evolving user standards with minimal retraining. Existing fine-tuning methods typically condition predictions on pre-defined policies, restricting their generalizability to new policies or necessitating extensive retraining to adapt. Conversely, training-free methods struggle with limited context lengths, making it difficult to incorporate all the policies comprehensively. To overcome these limitations, we propose to condition model's judgment on "precedents", which are the reasoning processes of prior data points similar to the given input. By leveraging precedents instead of fixed policies, our approach greatly enhances the flexibility and adaptability of the guardrail. In this paper, we introduce a critique-revise mechanism for collecting high-quality precedents and two strategies that utilize precedents for robust prediction. Experimental results demonstrate that our approach outperforms previous methods across both few-shot and full-dataset scenarios and exhibits superior generalization to novel policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。