提升大模型逻辑判断一致性,让其决策更稳定可靠。
Aligning with Logic: Measuring, Evaluating and Improving Logical Preference Consistency in Large Language Models
- 基于传递性、交换性和否定不变性构建评估框架。
- 发现现有大模型在逻辑判断上存在明显不一致现象。
- 提出REPAIR方法,在保持人类偏好对齐的同时增强逻辑一致性。
大语言模型(LLMs)需具备可预测性和可信度以支持可靠的决策系统。然而当前模型常表现出判断不一致。本文将逻辑偏好一致性视为构建更可靠LLM系统的基础要求,确保决策的稳定与连贯,减少随机或矛盾输出。为量化逻辑偏好一致性,我们提出一个基于传递性、交换性和否定不变性三个基本性质的通用评估框架。在多种大模型上的广泛实验表明,这些性质是判断鲁棒性的强指标。此外,我们引入数据精炼与增强技术REPAIR,可在保持与人类偏好对齐的同时提升逻辑一致性。最后,我们证明提升一致性能显著改善基于逻辑的大模型算法性能,强化决策系统的稳定性与连贯性。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are expected to be predictable and trustworthy to support reliable decision-making systems. Yet current LLMs often show inconsistencies in their judgments. In this work, we examine logical preference consistency as a foundational requirement for building more dependable LLM systems, ensuring stable and coherent decision-making while minimizing erratic or contradictory outputs. To quantify the logical preference consistency, we propose a universal evaluation framework based on three fundamental properties: transitivity, commutativity and negation invariance. Through extensive experimentation across diverse LLMs, we demonstrate that these properties serve as strong indicators of judgment robustness. Furthermore, we introduce a data refinement and augmentation technique, REPAIR, that enhances logical consistency while maintaining alignment with human preferences. Finally, we show that improving consistency leads to better performance in LLM-driven logic-based algorithms, reinforcing stability and coherence in decision-making systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。