提升抽取式问答模型鲁棒性,解决训练数据缺陷问题。
Towards Robust Extractive Question Answering Models: Rethinking the Training Methodology
- 设计新损失函数,打破传统数据集隐含假设。
- 跨域测试F1提升5.7分,对抗攻击下性能下降仅三分之一。
- 适合关注模型可靠性与真实场景应用的研究者。
本文提出一种新型训练方法以提升抽取式问答(EQA)模型的鲁棒性。以往研究表明,现有模型在包含不可回答问题的EQA数据集上训练时,对分布偏移和对抗攻击表现出显著脆弱性。尽管如此,将不可回答问题纳入训练数据对保障实际应用可靠性至关重要。本文提出的训练方法包含针对EQA问题的新损失函数,并挑战了多个EQA数据集中存在的隐含假设。使用该方法训练的模型在保持域内性能的同时,在域外数据集上实现显著提升,所有测试集平均F1得分提高5.7。此外,模型在两种对抗攻击下表现更稳健,性能下降幅度仅为基准模型的约三分之一。
原文摘要 · Abstract (English)
This paper proposes a novel training method to improve the robustness of Extractive Question Answering (EQA) models. Previous research has shown that existing models, when trained on EQA datasets that include unanswerable questions, demonstrate a significant lack of robustness against distribution shifts and adversarial attacks. Despite this, the inclusion of unanswerable questions in EQA training datasets is essential for ensuring real-world reliability. Our proposed training method includes a novel loss function for the EQA problem and challenges an implicit assumption present in numerous EQA datasets. Models trained with our method maintain in-domain performance while achieving a notable improvement on out-of-domain datasets. This results in an overall F1 score improvement of 5.7 across all testing sets. Furthermore, our models exhibit significantly enhanced robustness against two types of adversarial attacks, with a performance decrease of only about a third compared to the default models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。