arXiv:2601.17481cs.AI2026-01

自构建动态防护网,让对话模型实时抵御新威胁

Lattice: Generative Guardrails for Conversational Agents

  • 通过模拟与优化自动构建初始防护规则
  • 在跨领域数据上实现7%的准确率提升
  • 适合需要持续安全升级的对话系统开发者

对话式AI系统需防范有害输出,但现有方法依赖静态规则,难以应对新威胁或部署环境变化。我们提出Lattice框架,可自主构建并持续优化防护机制。该框架分两阶段运行:构建阶段通过迭代模拟与优化,从标注样本生成初始防护规则;持续改进阶段则通过风险评估、对抗测试和规则整合,自动适应已部署系统。在ProsocialDialog数据集上的评估显示,Lattice在保留数据上达到91%的F1分数,优于关键词基线43个百分点,优于LlamaGuard 25个百分点,优于NeMo 4个百分点。持续改进阶段在跨域数据上实现7个百分点的F1提升,得益于闭环优化。实验表明,有效防护规则可通过迭代优化实现自生成。

原文摘要 · Abstract (English)

Conversational AI systems require guardrails to prevent harmful outputs, yet existing approaches use static rules that cannot adapt to new threats or deployment contexts. We introduce Lattice, a framework for self-constructing and continuously improving guardrails. Lattice operates in two stages: construction builds initial guardrails from labeled examples through iterative simulation and optimization; continuous improvement autonomously adapts deployed guardrails through risk assessment, adversarial testing, and consolidation. Evaluated on the ProsocialDialog dataset, Lattice achieves 91% F1 on held-out data, outperforming keyword baselines by 43pp, LlamaGuard by 25pp, and NeMo by 4pp. The continuous improvement stage achieves 7pp F1 improvement on cross-domain data through closed-loop optimization. Our framework shows that effective guardrails can be self-constructed through iterative optimization.

对话安全自适应防护持续学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。