用程序生成+统计校准,让大模型推理更可靠。
An Empirical Study of Conformal Prediction in LLM with ASP Scaffolds for Robust Reasoning
- 让大模型生成ASP程序集,用校准方法保证输出正确性
- 在复杂推理任务上准确率显著提升,最高达87%
- 适合需要高可靠性推理的场景,如数学证明、逻辑推演
本文研究将共形语言建模(CLM)与答案集编程(ASP)结合,提升标准开放权重大模型在复杂多步推理任务中的表现。基于需空间推理的StepGame数据集,利用CLM从大模型生成一组ASP程序,为输出提供统计正确性保障。实验表明,CLM显著优于使用标准采样方法的基线模型,在不同推理复杂度下均实现明显准确率提升。此外,使用大模型作为评判者(LLM-as-Judge)可进一步优化CLM性能,尤其在评估结构与逻辑正确的ASP输出方面。然而,采用多样化校准集进行校准并未提升长链推理任务的泛化能力,表明其在处理更复杂任务时存在局限。
原文摘要 · Abstract (English)
In this paper, we examine the use of Conformal Language Modelling (CLM) alongside Answer Set Programming (ASP) to enhance the performance of standard open-weight LLMs on complex multi-step reasoning tasks. Using the StepGame dataset, which requires spatial reasoning, we apply CLM to generate sets of ASP programs from an LLM, providing statistical guarantees on the correctness of the outputs. Experimental results show that CLM significantly outperforms baseline models that use standard sampling methods, achieving substantial accuracy improvements across different levels of reasoning complexity. Additionally, the LLM-as-Judge metric enhances CLM's performance, especially in assessing structurally and logically correct ASP outputs. However, calibrating CLM with diverse calibration sets did not improve generalizability for tasks requiring much longer reasoning steps, indicating limitations in handling more complex tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。