arXiv:2604.16565cs.LGcs.AI2026-04中稿 · ICML被引 2

通过几何稳定性检测推理路径,让扩散语言模型自我验证答案正确性。

Reasoning on the Manifold: Bidirectional Consistency for Self-Verification in Diffusion Language Models

论文配图:Reasoning on the Manifold: Bidirectional Consistency for Self-Verification in Diffusion Language Models
图 1 · 摘自论文原文
  • 从流形角度建模推理轨迹,用双向一致性评估生成序列稳定性。
  • 无需标注即可判断解题有效性,支持复杂任务资源重分配。
  • 适用于诊断、推理优化与对齐训练,提升模型自纠错能力。

尽管扩散大语言模型(dLLMs)在全局规划上具有结构优势,但如何高效验证其生成的答案是否正确且推理过程合理仍是关键挑战。本文提出一种几何视角:在流形上进行推理。我们假设有效生成路径位于学习分布的高密度流形上,作为稳定吸引子;而无效路径则出现流形外漂移。为此,我们引入无需训练的无监督度量——双向流形一致性(BMC),通过前向掩码与反向重建循环量化生成序列的稳定性。实证表明,BMC在完整推理生命周期中表现出色:(1) 在诊断阶段,无需真实答案即可稳健区分解法有效性;(2) 在推理阶段,支持拒绝采样,将计算资源集中于复杂任务;(3) 在对齐阶段,作为密集几何奖励,将稀疏结果监督转化为细粒度指导,使模型实现超越基准的自我演化。结果确立了内在几何稳定性作为dLLMs正确性的可靠指标。

原文摘要 · Abstract (English)

While Diffusion Large Language Models (dLLMs) offer structural advantages for global planning, efficiently verifying that they arrive at correct answers via valid reasoning traces remains a critical challenge. In this work, we propose a geometric perspective: Reasoning on the Manifold. We hypothesize that valid generation trajectories reside as stable attractors on the high-density manifold of the learned distribution, whereas invalid paths exhibit off-manifold drift. To operationalize this, we introduce Bidirectional Manifold Consistency (BMC), a training-free, unsupervised metric that quantifies the stability of the generated sequence through a forward-masking and backward-reconstruction cycle. Empirically, we demonstrate BMC's versatility across the full reasoning lifecycle: (1) in Diagnosis, it serves as a robust discriminator of solution validity without ground truth answer; (2) in Inference, it enables rejection resampling to effectively concentrate computational resources on complex reasoning tasks; and (3) in Alignment, it functions as a dense geometric reward that transforms sparse outcome supervision into fine-grained guidance, empowering models to self-evolve beyond standard baselines. Our results establish intrinsic geometric stability as a robust indicator of correctness for dLLMs.

扩散模型推理验证自验证流形学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。