arXiv:2506.08123cs.CL2025-06EMNLP被引 26

用可解释的问答评估提升大模型对齐效果

QA-LIGN: Aligning LLMs through Constitutionally Decomposed QA

  • 将单一奖励拆解为具体原则的问答评估,实现透明反馈
  • 攻击成功率降低68.7%,误拒率仅0.67%,安全与有用性平衡更优
  • 适合关注模型安全、可解释训练的开发者和研究者

大语言模型(LLMs)在对齐如帮助性、诚实性和无害性等原则时,通常依赖于难以理解的标量奖励信号。我们提出QA-LIGN,通过结构化的自然语言程序将单一奖励分解为可解释的原则特异性评估。模型采用草稿-批评-修订流程,在GRPO训练中,基于评分标准进行符号化评估,为初始与修订后的响应提供透明反馈。应用于未过滤的Llama-3.1-8B-Instruct模型,QA-LIGN将攻击成功率降低高达68.7%,同时保持0.67%的假拒绝率,实现了帕累托最优的安全-有用性表现,并在同等训练条件下优于DPO和GRPO结合先进奖励模型的效果。结果表明,使奖励信号可解释且模块化能显著提升对齐效果,提示透明性有助于增强大模型安全性。

原文摘要 · Abstract (English)

Alignment of large language models (LLMs) with principles like helpfulness, honesty, and harmlessness typically relies on scalar rewards that obscure which objectives drive the training signal. We introduce QA-LIGN, which decomposes monolithic rewards into interpretable principle-specific evaluations through structured natural language programs. Models learn through a draft, critique, and revise pipeline, where symbolic evaluation against the rubrics provides transparent feedback for both initial and revised responses during GRPO training. Applied to uncensored Llama-3.1-8B-Instruct, QA-LIGN reduces attack success rates by up to 68.7% while maintaining a 0.67% false refusal rate, achieving Pareto optimal safety-helpfulness performance and outperforming both DPO and GRPO with state-of-the-art reward models given equivalent training. These results demonstrate that making reward signals interpretable and modular improves alignment effectiveness, suggesting transparency enhances LLM safety.

模型对齐可解释性安全评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。