用逻辑公式控制数据增强,让大模型生成更多样且严谨的推理题。
LFC-DA: Logical Formula-Controlled Data Augmentation for Enhanced Logical Reasoning
- 基于命题逻辑构建规则库,系统搜索有效公式
- 在ReClor和LogiQA上提升预训练模型推理准确率
- 适合需要高质量逻辑推理数据的研究者
针对复杂逻辑数据增强中人工标注成本高、大模型直接生成结果不可解释且同质化的问题,本文提出符号逻辑控制的数据增强方法LFC-DA。该方法将自然语言映射为命题表达式,构建紧凑规则库,并通过有界状态空间搜索系统发现有效逻辑公式,再将其还原为自然语言问题,确保生成内容兼具多样性与逻辑严谨性。在ReClor和LogiQA数据集上的实验表明,LFC-DA显著提升了预训练模型的逻辑推理能力,验证了其在大模型驱动逻辑数据增强中的有效性。
原文摘要 · Abstract (English)
For complex logical data augmentation, heavy reliance on human annotation is costly, whereas direct generation with large language models yields uninterpretable and logically homogeneous examples. To address this, we present LFC-DA, a symbolic-logic-controlled pipeline: logical text is first mapped to propositional expressions, a compact rule library is compiled, and a bounded state-space search systematically discovers valid formulas that are then verbalized back into natural-language questions, ensuring both diversity and logical rigor under propositional logic. Experiments on ReClor and LogiQA show significant improvements in the logical-reasoning accuracy of pretrained models, confirming the effectiveness of LFC-DA for LLM-guided logical data augmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。