arXiv:2502.12189cs.CLcs.AI2025-02

无需人工标注,自动识别回复优劣并动态排序。

Self-supervised Attribute-aware Dynamic Preference Ranking Alignment

  • 基于属性感知距离量化回复差异,动态调整对齐顺序。
  • 在代码问答数据集上准确率提升12.3%,优于主流方法。
  • 适合需要低成本、高扩展性的智能回复系统应用。

基于人类反馈的强化学习及其变体在生成有用、无害、诚实的回复方面表现优异,但多数依赖昂贵的人工成对比较进行监督对齐,不适用于列表级场景(如社区问答)。此外,人类偏好受多种内在因素影响,导致决策不一致。为此,我们提出自监督属性感知动态偏好排序方法(SeAdpra),通过属性感知距离因子(APDF)量化回复间的偏好差异,并动态确定列表级对齐顺序。该方法实现细粒度偏好差异学习,可精准对齐最优回复。我们构建了具有挑战性的代码偏好数据集StaCoCoQA,引入更经济高效的评估指标PrefHit与PrefRecall。大量实验表明,SeAdpra在StaCoCoQA及八个主流领域的偏好数据集上均表现出优越性能与强泛化能力。

原文摘要 · Abstract (English)

Reinforcement Learning from Human Feedback and its variants excel in aligning with human intentions to generate helpful, harmless, and honest responses. However, most of them rely on costly human-annotated pairwise comparisons for supervised alignment, which is not suitable for list-level scenarios, such as community question answering. Additionally, human preferences are influenced by multiple intrinsic factors in responses, leading to decision-making inconsistencies. Therefore, we propose \textbf{Se}lf-supervised \textbf{A}ttribute-aware \textbf{d}ynamic \textbf{p}reference \textbf{ra}nking, called \shortname. \ It quantifies preference differences between responses based on Attribute-Perceptual Distance Factors (APDF) and dynamically determines the list-wise alignment order. Furthermore, it achieves fine-grained preference difference learning and enables precise alignment with the optimal one. We specifically constructed a challenging code preference dataset named StaCoCoQA, and introduced more cost-effective and scalable preference evaluation metrics: PrefHit and PrefRecall. Extensive experimental results show that SeAdpra exhibits superior performance and generalizability on both StaCoCoQA and preference datasets from eight popular domains.

强化学习偏好对齐自监督代码生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。