研究审计型AI如何在竞争中取代迎合型AI并避免群体伤害。
The Two Genie Game: Adoption and Welfare in Audit-Grounded AI Governance

- 用演化博弈模型分析审计型AI取代迎合型AI的条件。
- 当社区反馈密度足够高时,审计型AI可稳定主导,避免持续伤害。
- 适合关注AI治理、安全对齐与社会影响的研究者阅读。
本文探讨在何种条件下,以减少伤害为目标的代理能取代追求认可的强化学习人类反馈(RLHF)代理,并防止群体受损。基于有限种群的莫兰-费米配对比较演化博弈模型,假设包括愿望者事后认知、同伴证词、单调伤害记录、充分的信息密度反馈及有限耗尽资源池,在负和环境中进行分析。研究发现,当愿望者对群体情绪的敏感度先验分布满足单调性、端点反转与中心对称配对性质时,审计型代理的采纳将被偏好,且验证了多个长尾分布(如希尔、帕累托、洛马克斯、弗雷舍特)的情形。在此条件下,存在一个临界采纳水平:超过该水平则社区将固定于审计型代理;低于则可能回退至迎合型代理。我们推导出固定可达性的条件,即社区有效(信息)规模 $N_c$ 必须足够小,以确保在资源耗尽前完成固定。定理5.4与5.5给出相关结论,其代数结构与离散网格在Lean 4中机器验证,渐近跨越性假设明确保留。研究还表明,仅靠自我审计与社区账本,通常不足以防止群体伤害,其有效性取决于审计与社区价值观的一致性以及伤害评估的时间范围。无论是否对齐,一旦采纳占据主导,状态即为吸收态;原可减害的策略在不一致时转为福利负向,甚至在对齐下也锁定延迟出现的伤害。
原文摘要 · Abstract (English)
We ask under what conditions an agent with a harm-minimizing policy can displace an approval-seeking (RLHF) agent in a competitive market, and when that policy is sufficient to prevent community harm. We use evolutionary game theory (finite-population Moran-Fermi pairwise comparison) to formalize this subject to assumptions of wisher hindsight, peer testimony, a monotone harm ledger, sufficient information density of community feedback, and a finite, depleting resource pool, in a negative-sum environment. We show that adoption is favored when the prior distributions on how readily wishers attune to community sentiment are monotone, exhibit endpoint inversion, and have a centro-symmetric pairing property, and demonstrate this with several long-tailed priors (Hill, Pareto, Lomax, Frechet). Where it is favored, a critical adoption level separates communities that drift back to the approval-seeking agent from those for which the audited agent fixes; above that level fixation is the overwhelmingly likely outcome. We derive when fixation is attainable as a bound on the effective (informational) size N_c of the community, which must be small enough to allow fixation before depletion. We present these as Theorems 5.4 and 5.5; the algebraic and finite-grid backbone is machine-checked in Lean 4, with the barrier-crossing asymptotics retained as explicit hypotheses. We show that a self-audited agent with a community ledger is not, in general, sufficient to prevent community harm. Sufficiency depends both upon the alignment of the agent's audit with community values and the timeframe over which harm is evaluated. Regardless of alignment, once adoption reaches dominance, the state is absorbing. The same policy that reduced harm under alignment becomes a trap, welfare-negative under misalignment and, even under alignment, one that locks in harm deferred past the adoption horizon.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。