研究AI安全设计中规则与性格的平衡,发现性格可靠性比部署规模更重要。
Rules or Character? Scaling Laws for AI Safety Design

- 用资源分配比例建模规则与性格的安全策略,考虑规模带来的失效风险。
- 性格脆弱性是影响最优设计的核心因素,其变化可导致策略偏移0.50。
- 大规模下规则和性格的最优选择趋于一致,适合关注安全架构的决策者。
人工智能安全系统结合训练时的性格塑造(如基于人类反馈的强化学习、宪法型AI)与推理时的规则执行(如输出过滤器、安全分类器),但缺乏对其在部署规模扩大时最优平衡的正式分析。本文构建一个简化对比静态模型,将安全设计参数化为[0,1]区间内的资源分配比例α,包含规模相关的过滤器退化、共模故障及性格脆弱性——即塑造行为在新情境下退化或崩溃的风险。在乘法帕累托损伤模型下,推导出期望伤害的闭式解,并通过蒙特卡洛模拟进行尾部风险(CVaR)分析。在三种情景(乐观、中等、悲观)下,最优α*位于内部或仅依赖规则边界,随部署规模T增长而轻微至显著向性格塑造倾斜,Δα*从+0.01到+0.21不等。主导参数为基线性格脆弱率p^(0)_frag,其变化可使α*偏移0.50,远超尾部严重性、过滤器质量或共模故障概率的影响。大尺度下CVaR与期望伤害最优解趋同。结果表明,安全架构决策更取决于性格塑造在分布偏移下的可靠性,而非部署规模本身。
原文摘要 · Abstract (English)
Artificial Intelligence (AI) safety systems combine character shaping (e.g., Reinforcement Learning from Human Feedback [RLHF], Constitutional AI), which modifies behavioral distributions at training time, with rule enforcement (e.g., output filters, safety classifiers), which blocks harmful outputs at inference time, yet little formal analysis exists on how their optimal balance should change as deployment scales increase. We introduce a stylized comparative-statics model that parameterizes safety design as a resource allocation alpha in [0,1] between these two approaches, incorporating scale-dependent filter degradation, common-mode failures, and character fragility -- the risk that shaped behavior degrades or collapses under novel conditions. Under a multiplicative Pareto damage model, we derive closed-form expected harm and supplement it with tail-risk (CVaR) analysis via Monte Carlo simulation. Across three scenarios (optimistic, moderate, pessimistic), the optimal alpha* is interior or at the rules-only boundary and shifts weakly toward character shaping as deployment scale T grows, from negligible (Delta alpha* = +0.01) to pronounced (Delta alpha* = +0.21) depending on scenario. The dominant parameter is the baseline character fragility rate p^(0)_frag, which shifts alpha* by 0.50 across its range -- far exceeding the effect of tail severity, filter quality, or common-mode failure probability. CVaR and expected-harm optima converge at large T. These results suggest that safety architecture decisions depend less on deployment scale per se than on the reliability of character shaping under distributional shift.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。