用少量标注数据让模型生成靠谱解释,且在新场景下表现稳定。
Self-Rationalization in the Wild: A Large Scale Out-of-Distribution Evaluation on NLI-related tasks
- 用少量标注数据微调大模型,实现跨领域解释生成。
- 仅需少量样本即可提升模型在19个新数据集上的解释能力。
- 解释质量与预测准确率正相关,人类评价最认可该指标。
自由文本解释表达力强、易理解,但多数数据集缺乏标注解释,制约可解释性模型训练。为此,我们研究如何利用已有解释数据集进行自理性生成,并评估模型在分布外(OOD)任务中的表现。对T5-Large和OLMo-7B模型进行微调,考察数据质量、微调样本数量及少样本选择方法的影响。在自然语言推理(NLI)、事实核查和摘要幻觉检测三类任务上,使用19个多样化OOD数据集进行评估。针对生成解释的评价,开展13个模型的人工评测,分析其与Acceptability分数(T5-11B)及其他三种基于LLM的无参考指标的相关性。结果表明:人类评价中,Acceptability分数与判断最相关,证明其有效性;1)少量标注样本即可有效适配模型用于分布外解释生成;2)微调数据来源比样本选择策略对分布外性能影响更大;3)预测准确率更高的模型往往产生更优解释,表现为更高接受度评分。
原文摘要 · Abstract (English)
Free-text explanations are expressive and easy to understand, but many datasets lack annotated explanation data, making it challenging to train models for explainable predictions. To address this, we investigate how to use existing explanation datasets for self-rationalization and evaluate models' out-of-distribution (OOD) performance. We fine-tune T5-Large and OLMo-7B models and assess the impact of fine-tuning data quality, the number of fine-tuning samples, and few-shot selection methods. The models are evaluated on 19 diverse OOD datasets across three tasks: natural language inference (NLI), fact-checking, and hallucination detection in abstractive summarization. For the generated explanation evaluation, we conduct a human study on 13 selected models and study its correlation with the Acceptability score (T5-11B) and three other LLM-based reference-free metrics. Human evaluation shows that the Acceptability score correlates most strongly with human judgments, demonstrating its effectiveness in evaluating free-text explanations. Our findings reveal: 1) few annotated examples effectively adapt models for OOD explanation generation; 2) compared to sample selection strategies, fine-tuning data source has a larger impact on OOD performance; and 3) models with higher label prediction accuracy tend to produce better explanations, as reflected by higher Acceptability scores.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。