arXiv:2510.06096cs.LGcs.CL2025-10被引 1

用贝叶斯方法让大模型目标可审计,解决对齐模糊性问题。

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives

  • 基于贝叶斯逆强化学习,推断模型目标的分布而非单一值。
  • 实证显示能逐步减少目标不确定性,提升对齐可靠性。
  • 适合安全团队和监管者用于验证模型真实意图。

大型语言模型(LLM)隐含优化的目标仍高度不透明,导致可信对齐与审计成为重大挑战。尽管逆强化学习(IRL)可从行为中推断奖励函数,但现有方法要么给出单一、过度自信的估计,要么无法解决任务的根本模糊性(非可辨识性)。本文提出一种严谨的审计框架,将奖励推断从简单估计重构为全面验证过程。该框架利用贝叶斯IRL不仅恢复目标的分布,还实现三项关键审计能力:(i) 通过多轮证据的后验收缩,量化并系统降低非可辨识性;(ii) 提供具有不确定性的可行动诊断,揭示虚假捷径,并识别分布外提示下不可信的推理;(iii) 验证策略级效用,证明经过精炼的低不确定性奖励可直接用于RLHF,实现与真实对齐过程相当的训练动态和毒性降低效果。实验上,该框架成功审计了一款净化后的LLM,获得校准良好且可解释的目标,增强了对齐保证。整体而言,本工作为审计员、安全团队与监管机构提供了一套实用工具,以验证大模型的真实目标,推动更可信、可问责的人工智能发展。

原文摘要 · Abstract (English)

The objectives that Large Language Models (LLMs) implicitly optimize remain dangerously opaque, making trustworthy alignment and auditing a grand challenge. While Inverse Reinforcement Learning (IRL) can infer reward functions from behaviour, existing approaches either produce a single, overconfident reward estimate or fail to address the fundamental ambiguity of the task (non-identifiability). This paper introduces a principled auditing framework that re-frames reward inference from a simple estimation task to a comprehensive process for verification. Our framework leverages Bayesian IRL to not only recover a distribution over objectives but to enable three critical audit capabilities: (i) Quantifying and systematically reducing non-identifiability by demonstrating posterior contraction over sequential rounds of evidence; (ii) Providing actionable, uncertainty-aware diagnostics that expose spurious shortcuts and identify out-of-distribution prompts where the inferred objective cannot be trusted; and (iii) Validating policy-level utility by showing that the refined, low-uncertainty reward can be used directly in RLHF to achieve training dynamics and toxicity reductions comparable to the ground-truth alignment process. Empirically, our framework successfully audits a detoxified LLM, yielding a well-calibrated and interpretable objective that strengthens alignment guarantees. Overall, this work provides a practical toolkit for auditors, safety teams, and regulators to verify what LLMs are truly trying to achieve, moving us toward more trustworthy and accountable AI.

大模型对齐贝叶斯方法安全审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。