通过群体比较提升大模型推理一致性,兼顾多样性。
DGPO: Beyond Pairwise Preferences with Directional Consistent Groupwise Optimization

- 用多候选对比构建群体级监督信号,显式建模方向一致性。
- 在五个基准上平均提升3.2%,部分场景达3.6%准确率增益。
- 适合需要可靠推理路径的复杂任务,如数学与逻辑推理。
尽管大型语言模型(LLMs)取得了显著进展,当前的偏好优化方法仍难以在保持推理多样性的同时实现方向一致性对齐。为此,我们提出方向性群体偏好优化(DGPO),一种轻量级框架,通过在群体层面聚合监督信号,并利用多候选对比显式建模方向感知的一致性。DGPO将正向与反向问答实例组织成结构化集合,优化基于间隔的概率目标,以区分连贯的推理路径与不一致的替代路径。该群体形式捕捉比成对目标更丰富的相对信息,并强化多样推理路径间的一致性。实证结果表明,我们构建的反向数据在五个基准上带来3.2%的平均提升,而DGPO在多个数据集和模型家族中均实现持续增益,最高达成3.6%的平均准确率提升。
原文摘要 · Abstract (English)
Although Large Language Models (LLMs) have made remarkable progress, current preference optimization methods still struggle to align directional consistency while preserving reasoning diversity. To address this limitation, we propose Directional-Groupwise Preference Optimization (DGPO), a lightweight framework that aggregates supervision signals at the group level and explicitly models direction-aware alignment through multi-candidate comparisons. DGPO organizes forward and reverse question-answer instances into structured sets and optimizes a margin-based likelihood objective that separates coherent reasoning paths from inconsistent alternatives. This group-wise formulation captures richer relative information than pairwise objectives and reinforces consistency across diverse reasoning pathways. Empirical results show that our constructed reverse data yields a 3.2% average improvement across five benchmarks, while DGPO further delivers consistent gains across multiple datasets and model families, achieving average accuracy improvements of up to 3.6%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。