量化NLP模型预测不确定性,提升可靠性与可信度。
On Uncertainty In Natural Language Processing
- 从语言、统计与神经视角分析NLP中的不确定性来源。
- 跨三种语言的实验证明方法可有效降低预测不确定性。
- 适合关注模型可信度与安全部署的研究者与开发者。
过去十年深度学习的发展催生了日益强大的系统,并广泛应用于各类场景。在自然语言处理领域,大语言模型的突破推动了诸多面向用户的应用。为充分受益于该技术并减少潜在危害,量化模型预测的可靠性及其背后的不确定性至关重要。本文从语言学、统计学和神经网络角度研究了自然语言处理中的不确定性特征,并通过实验设计优化来减少和量化不确定性。研究进一步理论与实证分析了文本分类任务中归纳偏置对不确定性的影响,涵盖丹麦语、英语、芬兰语三种语言及多个任务,对比了大量不确定性量化方法。此外,提出一种基于非交换型共形预测的校准采样方法,生成更紧致且覆盖更准确的词元集合。最后,开发了一种仅依赖目标模型输入与输出文本的辅助预测器,用于量化大型黑箱语言模型的置信度。
原文摘要 · Abstract (English)
The last decade in deep learning has brought on increasingly capable systems that are deployed on a wide variety of applications. In natural language processing, the field has been transformed by a number of breakthroughs including large language models, which are used in increasingly many user-facing applications. In order to reap the benefits of this technology and reduce potential harms, it is important to quantify the reliability of model predictions and the uncertainties that shroud their development. This thesis studies how uncertainty in natural language processing can be characterized from a linguistic, statistical and neural perspective, and how it can be reduced and quantified through the design of the experimental pipeline. We further explore uncertainty quantification in modeling by theoretically and empirically investigating the effect of inductive model biases in text classification tasks. The corresponding experiments include data for three different languages (Danish, English and Finnish) and tasks as well as a large set of different uncertainty quantification approaches. Additionally, we propose a method for calibrated sampling in natural language generation based on non-exchangeable conformal prediction, which provides tighter token sets with better coverage of the actual continuation. Lastly, we develop an approach to quantify confidence in large black-box language models using auxiliary predictors, where the confidence is predicted from the input to and generated output text of the target model alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。