用用户自定义政策训练可扩展的过滤器,兼顾安全与多样合规需求。
PAM: Training Policy-Aligned Moderation Filters at Scale
- 基于用户政策自动生成训练数据,无需人工标注
- 在多个政策基准上超越现有模型,且推理速度提升5-100倍
- 适合需定制化合规策略的应用场景,如医疗、饮食、文化等
大型语言模型仍易出现对齐偏差和越狱攻击,外部防护如内容过滤器至关重要。现有过滤器多聚焦于安全,难以满足实际部署中的多样化对齐需求。本文提出政策对齐过滤(PAM),一种基于用户自定义政策的灵活框架,可训练特定应用场景的过滤器。PAM通过自动化生成训练数据,无需依赖人工编写样本,支持大规模、多样的对齐目标。其训练出的过滤器性能达到当前顶尖安全过滤器水平,并在四个新提出的用户标注政策执行基准(PAMbench)——涵盖年龄限制、饮食要求、文化适配和医疗建议限制——上表现更优。同时,相比条件推理模型,PAM过滤器推理速度提升5-100倍。
原文摘要 · Abstract (English)
Large language models (LLMs) remain vulnerable to misalignment and jailbreaks, making external safeguards like moderation filters essential, yet existing filters often focus narrowly on safety, falling short of the broader alignment needs seen in real-world deployments. We introduce Policy Aligned Moderation (PAM), a flexible framework for training custom moderation filters grounded in user-defined policies that extend beyond conventional safety objectives. PAM automates training data generation without relying on human-written examples, enabling scalable support for diverse, application-specific alignment goals and generation policies. PAM-trained filters match the performance of state-of-the-art safety moderation filters and policy reasoning models, and outperform them on PAMbench, four newly introduced user-annotated policy enforcement benchmarks that target age restrictions, dietary accommodations, cultural alignment, and limitations in medical guidance. These performance gains are achieved while the PAM filter runs 5-100x faster at inference than policy-conditioned reasoning models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。