研究大模型在少样本下如何检索和理解文档证据,发现检索错误是主要瓶颈。
Retrieving Versus Understanding Extractive Evidence in Few-Shot Learning
- 对比模型预测与人工标注证据的匹配度,分析检索与理解的关系。
- 五数据集实验显示预测错误与检索错误高度相关,但与理解错误关联弱。
- 结果表明改进检索机制可显著提升下游任务表现,适合模型优化方向参考。
对大语言模型在少样本设置下使用文档内证据进行文档级决策的能力进行分析。我们通过两个主流闭源模型,在五个数据集上评估模型预测错误与黄金标准人工标注提取证据之间的检索错误关联性。通过两次消融实验,研究当标签预测和证据检索错误均能归因于相关证据质量的情况。结果表明,模型预测错误与证据检索错误存在强烈的经验关联,但证据检索错误大多不与证据解释错误相关,这对基于此机制的下游应用是一个积极信号。
原文摘要 · Abstract (English)
A key aspect of alignment is the proper use of within-document evidence to construct document-level decisions. We analyze the relationship between the retrieval and interpretation of within-document evidence for large language model in a few-shot setting. Specifically, we measure the extent to which model prediction errors are associated with evidence retrieval errors with respect to gold-standard human-annotated extractive evidence for five datasets, using two popular closed proprietary models. We perform two ablation studies to investigate when both label prediction and evidence retrieval errors can be attributed to qualities of the relevant evidence. We find that there is a strong empirical relationship between model prediction and evidence retrieval error, but that evidence retrieval error is mostly not associated with evidence interpretation error--a hopeful sign for downstream applications built on this mechanism.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。