通过反事实反馈提升大模型的因果推理能力
Reasoning Elicitation in Language Models via Counterfactual Feedback
- 设计新评估指标,同时衡量事实与反事实问题回答准确率
- 提出微调方法,使模型在归纳和演绎推理中表现更优
- 在真实场景验证,显著提升模型泛化能力
尽管语言模型日益有效,其推理能力仍不充分,尤其在反事实问答中的因果推理能力薄弱。本文首先提出新的评估指标,平衡事实与反事实问题的准确率,更全面地刻画模型推理能力;其次,设计多种微调方法以激发更好的推理机制;最后,在多种真实场景下评估微调后模型的表现。实验表明,所提方法在需要归纳与演绎推理的问题上,系统性提升了基线模型的泛化能力。
原文摘要 · Abstract (English)
Despite the increasing effectiveness of language models, their reasoning capabilities remain underdeveloped. In particular, causal reasoning through counterfactual question answering is lacking. This work aims to bridge this gap. We first derive novel metrics that balance accuracy in factual and counterfactual questions, capturing a more complete view of the reasoning abilities of language models than traditional factual-only based metrics. Second, we propose several fine-tuning approaches that aim to elicit better reasoning mechanisms, in the sense of the proposed metrics. Finally, we evaluate the performance of the fine-tuned language models in a variety of realistic scenarios. In particular, we investigate to what extent our fine-tuning approaches systemically achieve better generalization with respect to the base models in several problems that require, among others, inductive and deductive reasoning capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。