用逻辑命题构建测试集,评估大模型的推理能力。
Rosetta-PL: Propositional Logic as a Benchmark for Large Language Model Reasoning
- 将Lean中的逻辑命题转为自定义语言,用于微调大模型。
- 超过2万样本后准确率趋于稳定,保持逻辑关系能显著提升精度。
- 适合研究形式化推理与低资源语言任务的学者参考。
大型语言模型(LLMs)主要在高资源自然语言上训练,限制了其在低资源环境和需要深度逻辑推理任务中的表现。本研究提出Rosetta-PL,一个用于评估LLMs在受控环境中逻辑推理与泛化能力的基准。通过将来自Lean的数据集中的逻辑命题翻译成一种自定义逻辑语言,并用于微调大模型(如GPT-4o),我们分析了数据集规模和翻译方法对模型性能的影响。结果表明,在翻译过程中保持逻辑关系可显著提升精确度,且准确率在约20,000个训练样本后趋于平稳。这些发现为优化大模型在形式化推理任务中的训练提供了重要指导,并有助于提升其在多种低资源语言应用中的表现。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are primarily trained on high-resource natural languages, limiting their effectiveness in low-resource settings and in tasks requiring deep logical reasoning. This research introduces Rosetta-PL, a benchmark designed to evaluate LLMs' logical reasoning and generalization capabilities in a controlled environment. We construct Rosetta-PL by translating a dataset of logical propositions from Lean into a custom logical language, which is then used to fine-tune an LLM (e.g., GPT-4o). Our experiments analyze the impact of the size of the dataset and the translation methodology on the performance of the model. Our results indicate that preserving logical relationships in the translation process significantly boosts precision, with accuracy plateauing beyond roughly 20,000 training samples. These insights provide valuable guidelines for optimizing LLM training in formal reasoning tasks and improving performance in various low-resource language applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。