用安全策略动态调控新策略,实现零风险探索与性能提升。
Conformal Policy Control
- 以安全策略数据进行共形校准,动态设定新策略行动强度。
- 首次在非单调损失下提供有限样本保证,支持高风险场景部署。
- 适用自然语言到生物分子设计,适合追求安全高效探索的研究者。
智能体需尝试新行为以探索和优化。但在高风险环境中,违反安全约束可能导致伤害,迫使智能体下线,中断后续交互。模仿旧行为虽安全,但过度保守会抑制探索。行为改变多少才算过界?本文提出一种方法:利用任意安全参考策略作为概率调节器,对任意未经测试的优化策略进行控制。基于安全策略的数据进行共形校准,可确定新策略可采取的激进程度,同时严格遵守用户声明的风险容忍度。与传统保守优化不同,本方法无需假设用户已识别正确模型类别或调优超参数。与以往共形方法相比,理论首次提供非单调有界损失函数下的有限样本保证,并引入新的策略控制范式。实验覆盖从自然语言问答到生物分子工程的应用,证明从首次部署起即可实现安全探索,并能提升性能。
原文摘要 · Abstract (English)
An agent must try new behaviors to explore and improve. In high-stakes environments, an agent that violates safety constraints may cause harm and must be taken offline, curtailing any future interaction. Imitating old behavior is safe, but excessive conservatism discourages exploration. How much behavior change is too much? We show how to use any safe reference policy as a probabilistic regulator for any optimized but untested policy. Conformal calibration on data from the safe policy determines how aggressively the new policy can act, while provably enforcing the user's declared risk tolerance. Unlike conservative optimization methods, we do not assume the user has identified the correct model class nor tuned any hyperparameters. Unlike previous conformal methods, our theory provides finite-sample guarantees even for non-monotonic bounded loss functions, and it introduces a new policy control setting. Our experiments on applications ranging from natural language question answering to biomolecular engineering show that safe exploration is not only possible from the first moment of deployment, but can also improve performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。