arXiv:2510.03485cs.AI2025-10中稿 · ICML被引 2

轻量级模型实现高效合规检测,支持长任务中的政策遵循。

Learning Efficient Guardrails for Compliance

  • 构建6万条策略轨迹数据集,支持完整与前缀式违规检测。
  • 轻量模型在多领域上保持高检测准确率与推理效率。
  • 适合关注长周期任务合规性的研究者与开发者使用。

自主网络代理在执行长周期任务时日益普及,但其对现实世界政策的遵守能力仍远落后于标准安全目标。为填补这一空白,我们提出PolicyGuardBench,一个包含60,000条策略-轨迹对的数据集,用于评估全轨迹与新型前缀式违规检测任务。基于该数据集,我们训练了PolicyGuard——一种轻量级守卫模型,在保持高推理效率的同时实现强检测精度。值得注意的是,该模型展现出稳健的泛化能力,即使在未见领域仍能维持高性能。这些贡献建立了一个全面的研究框架,证明在小规模下实现准确且可泛化的守卫机制是可行的。

原文摘要 · Abstract (English)

Autonomous web agents are increasingly deployed for long-horizon tasks, yet their ability to adhere to real-world policies remains critically underexplored compared to standard safety objectives. To address this gap, we introduce PolicyGuardBench, a benchmark of 60k policy-trajectory pairs designed to evaluate compliance through both full-trajectory and novel prefix-based violation detection tasks. Using this dataset, we train PolicyGuard, a lightweight guardrail model that achieves strong detection accuracy while maintaining high inference efficiency. Notably, our model demonstrates robust generalization capabilities, preserving high performance even on unseen domains. These contributions establish a comprehensive framework for studying policy compliance, showing that accurate and generalizable guardrails are feasible at small scales.

合规检测轻量模型长周期任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。