对比不同自然语言解释对上下文学习的影响,发现外部生成解释效果最佳。
When Do Explanations Help In-Context Learning? A Comparative Study of Natural Language Explanation Types and Faithfulness

- 比较人类、自动生成和外部LLM生成的解释在上下文学习中的表现
- 外部生成解释在分类任务中显著提升准确率,接近人工解释效果
- 解释的忠实度筛选对模型性能影响因任务和指标而异
自然语言解释(NLEs)被越来越多地用作输入,例如作为少量示例推理来影响上下文学习(ICL)中的模型行为。然而,不同类型的NLEs在解释增强提示下的下游性能影响仍不明确。为此,我们在六个基准上对四款指令微调模型进行了对比评估,研究NLE来源(人类撰写、自生成解释、外部LLM生成)和选择策略(随机选择与基于忠实度的过滤)对下游实用性的影。结果显示,在分类类基准上,添加NLE可提升准确率;其中外部生成的LLM-NLE在多数情况下表现优异,且在有可用时与人工解释相当,而自生成解释更依赖选择策略。在数学推理任务中,效果则高度依赖模型和来源。基于忠实度的自生成解释筛选整体带来小幅平均提升,但可能提高或降低性能,取决于任务、指标和模型。不同忠实度度量结果差异显著,影响所选示例及其预测效用。随机替换和分布外解释的鲁棒性测试显示部分稳健性,表明语义对齐有助于性能提升。总体而言,研究为实际提示流程中解释的选择与报告提供了依据。
原文摘要 · Abstract (English)
Natural language explanations (NLEs) are increasingly used as inputs, for example, as few-shot rationales that influence model behavior in in-context learning (ICL). However, it remains unclear how different types of NLEs compare in their effects on downstream model performance in explanation-augmented prompting. Therefore, we provide a comparative evaluation across six benchmarks and four instruction-tuned models, studying how NLE source (human-written when available, self-generated explanations, generated by an external LLM) and NLE selection (random vs faithfulness-based filtering) affect downstream utility of NLEs when used in ICL settings. Our extensive evaluation shows that, on classification-style benchmarks, adding NLEs to few-shot prompts often improves accuracy over few-shot prompting without explanations; among NLE sources, externally generated LLM-NLEs often provide strong downstream utility and remain competitive with human rationales where both are available, whereas self-NLEs are more sensitive to the selection strategy. On math reasoning, the effects are more model- and source-dependent. We further show that faithfulness-based selection of self-NLEs yields small average gains overall, but can improve or reduce performance depending on the metric, task, and model. Different faithfulness metrics can disagree substantially, affecting which self-NLE examples are selected and their downstream predictive utility. Robustness tests with randomly swapped and out-of-distribution rationales indicate partial robustness, suggesting that semantic alignment contributes to performance gains. Overall, our results provide insights for selecting and reporting explanations that influence model behavior in practical prompting pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。