提出新评估框架HEAL,用多假设分析改进大模型对齐效果评测。
HEAL: A Hypothesis-Based Preference-Aware Analysis Framework
- 将对齐评估视为假设空间内的重排序过程,突破单响应局限。
- 发现现有方法能有效捕捉代理模型偏好并压制负样本。
- 适合研究对齐机制、优化算法的学者使用。
偏好优化方法如DPO在大模型对齐中表现优异,但现有评估仅依赖单一响应,忽略了真实场景中可能生成的其他输出。为此,本文提出一种基于假设的推理感知分析框架HEAL,将偏好对齐建模为假设空间内的重排序过程。该框架引入两个互补指标:排名准确率(评估序数一致性)与偏好强度相关性(评估连续对齐)。为支持该框架,我们构建了UniHypoBench——一个基于多样化指令-响应对的统一假设基准。通过大量实验,尤其聚焦偏好学习的内在机制,结果表明当前方法能有效捕获代理模型提供的偏好,同时抑制负样本。这一发现从理论和实践两方面推动了偏好学习研究:理论上,提出假设空间分析作为理解对齐的新范式;实践中,HEAL为优化方法提供强大诊断工具,并指明了更全面捕捉偏好的先进算法发展方向。
原文摘要 · Abstract (English)
Preference optimization methods like DPO have achieved remarkable performance in LLM alignment. However, the evaluation for these methods relies on a single response and overlooks other potential outputs, which could also be generated in real-world applications within this hypothetical space. To address this issue, this paper presents a \textbf{H}ypothesis-based Pr\textbf{E}ference-aware \textbf{A}na\textbf{L}ysis Framework (HEAL), a novel evaluation paradigm that formulates preference alignment as a re-ranking process within hypothesis spaces. The framework incorporates two complementary metrics: ranking accuracy for evaluating ordinal consistency and preference strength correlation for assessing continuous alignment. To facilitate this framework, we develop UniHypoBench, a unified hypothesis benchmark constructed from diverse instruction-response pairs. Through extensive experiments based on HEAL, with a particular focus on the intrinsic mechanisms of preference learning, we demonstrate that current preference learning methods can effectively capture preferences provided by proxy models while simultaneously suppressing negative samples. These findings contribute to preference learning research through two significant avenues. Theoretically, we introduce hypothesis space analysis as an innovative paradigm for understanding preference alignment. Practically, HEAL offers researchers robust diagnostic tools for refining preference optimization methods, while our empirical results identify promising directions for developing more advanced alignment algorithms capable of comprehensive preference capture.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。