arXiv:2505.23540cs.CL2025-05ACL被引 1

让大模型推理更可信:同时优化答案正确性和逻辑连贯性

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning

  • 用双重标准筛选偏好数据:答案对错 + 生成概率一致性
  • 在多个模型和测试集上超越纯结果导向方法
  • 适合追求推理严谨性的研究人员和应用开发

近期的偏好优化进展展示了显著提升大语言模型(LLMs)数学推理能力的潜力。现有方法虽利用高质量成对偏好数据,以答案正确性或一致性等结果指标为准则,但忽略了响应内部的逻辑连贯性。为此,我们提出概率一致性偏好优化(PCPO),建立双重定量指标用于偏好选择:(1) 表层答案正确性;(2) 响应中逐标记的概率一致性。大量实验表明,所提方法在多种模型与基准测试中均一致优于仅依赖结果指标的方法。代码已公开于 https://github.com/YunqiaoYang/PCPO。

原文摘要 · Abstract (English)

Recent advances in preference optimization have demonstrated significant potential for improving mathematical reasoning capabilities in large language models (LLMs). While current approaches leverage high-quality pairwise preference data through outcome-based criteria like answer correctness or consistency, they fundamentally neglect the internal logical coherence of responses. To overcome this, we propose Probability-Consistent Preference Optimization (PCPO), a novel framework that establishes dual quantitative metrics for preference selection: (1) surface-level answer correctness and (2) intrinsic token-level probability consistency across responses. Extensive experiments show that our PCPO consistently outperforms existing outcome-only criterion approaches across a diverse range of LLMs and benchmarks. Our code is publicly available at https://github.com/YunqiaoYang/PCPO.

大模型推理偏好优化逻辑一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。