用法律机制解决AI规则解释模糊问题,提升模型一致性。
Statutory Construction and Interpretation for Artificial Intelligence
- 借鉴法律制定与解释流程,设计规则优化与提示约束双机制。
- 在WildChat数据集5000场景中,判断一致性显著提升。
- 适合关注AI合规性与行为稳定性的研究者与开发者。
AI系统日益受自然语言规则约束,但其解释模糊性问题尚未得到充分重视:规则表述与应用均可能导致歧义。与法律体系通过上诉审查等制度控制歧义不同,当前AI对齐流程缺乏类似保障。同一规则的不同解释可能引发模型行为不一致或不稳定。本文基于法律理论,分析对齐流程在规则制定与应用阶段的模糊管理缺口,提出计算框架:(1) 规则精炼管道,通过修订模糊规则减少解释分歧(类比行政机关立法或迭代立法);(2) 基于提示的解释约束,降低规则应用中的不一致性(类比法律解释原则)。在WildChat数据集5000个场景的子集上评估,两项干预均显著提升合理解释者间的判断一致性。该方法为系统化应对解释模糊性迈出关键一步,是构建更鲁棒、守法的AI系统的必要进展。
原文摘要 · Abstract (English)
AI systems are increasingly governed by natural language principles, yet a key challenge arising from reliance on language remains underexplored: interpretive ambiguity. As in legal systems, ambiguity arises both from how these principles are written and how they are applied. But while legal systems use institutional safeguards to manage such ambiguity, such as transparent appellate review policing interpretive constraints, AI alignment pipelines offer no comparable protections. Different interpretations of the same rule can lead to inconsistent or unstable model behavior. Drawing on legal theory, we identify key gaps in current alignment pipelines by examining how legal systems constrain ambiguity at both the rule creation and rule application steps. We then propose a computational framework that mirrors two legal mechanisms: (1) a rule refinement pipeline that minimizes interpretive disagreement by revising ambiguous rules (analogous to agency rulemaking or iterative legislative action), and (2) prompt-based interpretive constraints that reduce inconsistency in rule application (analogous to legal canons that guide judicial discretion). We evaluate our framework on a 5,000-scenario subset of the WildChat dataset and show that both interventions significantly improve judgment consistency across a panel of reasonable interpreters. Our approach offers a first step toward systematically managing interpretive ambiguity, an essential step for building more robust, law-following AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。