arXiv:2509.20105cs.AI2025-09ICML

用量子思想提升大模型推理连贯性,让逻辑链条更清晰。

PEPS: Quantum-Inspired Reinforcement Learning for Coherent Reasoning Traces in LLMs

  • 借鉴量子态结构设计一致性奖励机制
  • 在GSM8K等数据集上显著提升推理连贯性
  • 适合需要严谨逻辑链的复杂推理任务

大语言模型在需要多步结构化逻辑的任务中常出现推理断链问题。本文提出一种量子启发式方法,将投影纠缠对态(PEPS)的保真度作为奖励信号,融入近端策略优化(PPO)框架,通过结构一致性引导学习,而非依赖直接监督或对比目标。该方法在涵盖算术、直觉和蕴涵推理的多个数据集(GSM8K、StrategyQA、EntailmentBank)上评估,采用多种连贯性指标进行验证。结果表明,该方法显著优于监督、对比及预训练基线模型,证明了量子启发保真度作为增强大模型推理连贯性的有效基础。

原文摘要 · Abstract (English)

Large Language Models (LLMs) often struggle with maintaining coherent multi-step reasoning traces, particularly in tasks that require a structured logical flow. This work introduces a quantum-inspired approach to address the challenge by incorporating a fidelity-based reward derived from Projected Entangled Pair States (PEPS) into Proximal Policy Optimization. Unlike prior approaches that use direct supervision or contrastive objectives, the proposed method guides learning through structural consistency, offering a novel approach to enforce global coherence in generated reasoning traces. The proposed framework is evaluated using multiple coherence-determining metrics on diverse datasets such as GSM8K, StrategyQA, and EntailmentBank spanning arithmetic, intuitive, and entailment-based reasoning. Results show that the proposed quantum-inspired approach offers significant improvements over supervised, contrastive, and pretrained baseline approaches, highlighting the effectiveness of quantum-inspired fidelity as a foundation to improve reasoning trace coherence in LLMs.

推理生成强化学习量子启发逻辑连贯

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。