提出新评估框架,检测奖励模型在真实扰动下的系统性弱点。
Reward Auditor: Inference on Reward Modeling Suitability in Real-World Perturbed Scenarios
- 通过假设检验框架审计奖励模型在真实扰动下的偏好置信度分布退化。
- 量化了奖励模型在多种真实场景中的脆弱性程度与统计显著性。
- 适合关注大模型对齐安全、鲁棒性的研究者和工程师使用。
可靠的奖励模型(RMs)对于确保大语言模型(LLMs)的安全对齐至关重要。然而,当前的RM评估方法仅关注特定场景下的偏好感知准确性,忽略了其在真实世界场景中的关键缺陷。我们发现真正的挑战在于评估一个新维度:适用性,即在特定现实扰动下的条件可靠性。为此,我们提出了Reward Auditor,一种专为RM适用性推断设计的假设检验框架。它不回答‘该RM对给定样本的偏好感知有多准确?’,而是通过科学审计回答:‘我们能否推断出该RM在特定真实场景中存在系统性漏洞?’ 在真实世界扰动场景下,Reward Auditor通过审计奖励模型偏好感知置信度的分布退化,量化统计显著性和效应大小。这使得能够在多样化真实场景中推断出奖励模型脆弱性的确定性和严重程度,为构建可验证安全、更鲁棒、更可信的下一代LLM对齐系统奠定基础。
原文摘要 · Abstract (English)
Reliable reward models (RMs) are critical for ensuring the safe alignment of large language models (LLMs). However, current RM evaluation methods focus solely on preference perception accuracies in given specific scenarios, obscuring the critical vulnerabilities of RMs in real-world scenarios. We identify the true challenge lies in assessing a novel dimension: Suitability, defined as conditional reliability under specific real-world perturbations. To this end, we introduce Reward Auditor, a hypothesis-testing framework specifically designed for RM suitability inference. Rather than answering "How accurate is the RM's preference perception for given samples?", it employs scientific auditing to answer: "Can we infer RMs exhibit systematic vulnerabilities in specific real-world scenarios?". Under real-world perturbed scenarios, Reward Auditor quantifies statistical significance and effect size by auditing distribution degradation of RM preference perception confidence. This enables inference of both the certainty and severity of RM vulnerabilities across diverse real-world scenarios. This lays a solid foundation for building next-generation LLM alignment systems that are verifiably safe, more robust, and trustworthy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。