用大模型生成更真实扰动,提升通用信息抽取的鲁棒性
Towards Robust Universal Information Extraction: Benchmark, Evaluation, and Solution
- 用大模型生成多样且自然的文本扰动,构建新基准数据集
- 仅用15%数据训练,模型在3个任务上平均提升7.5%性能
- 适合关注模型抗干扰能力与高效训练的研究者
本文旨在提升通用信息抽取(UIE)的鲁棒性,提出新基准数据集、全面评估方法及可行解决方案。现有鲁棒性基准存在两大局限:一是单一信息抽取任务扰动类型有限,难以有效评估UIE模型;二是依赖小模型或手工规则生成扰动,常导致对抗样本不自然。针对此,我们利用大语言模型(LLM)的强大生成能力,构建了名为RUIE-Bench的新鲁棒性基准数据集,涵盖多种信息抽取任务的多样化且真实的扰动。基于该数据集,我们全面评估现有UIE模型,发现基于LLM的模型及其他模型均出现显著性能下降。为提升鲁棒性并降低训练成本,我们提出一种动态选择困难样本进行迭代训练的数据增强方案。实验表明,仅使用15%的数据训练,即可在三个信息抽取任务上实现平均7.5%的相对性能提升。
原文摘要 · Abstract (English)
In this paper, we aim to enhance the robustness of Universal Information Extraction (UIE) by introducing a new benchmark dataset, a comprehensive evaluation, and a feasible solution. Existing robust benchmark datasets have two key limitations: 1) They generate only a limited range of perturbations for a single Information Extraction (IE) task, which fails to evaluate the robustness of UIE models effectively; 2) They rely on small models or handcrafted rules to generate perturbations, often resulting in unnatural adversarial examples. Considering the powerful generation capabilities of Large Language Models (LLMs), we introduce a new benchmark dataset for Robust UIE, called RUIE-Bench, which utilizes LLMs to generate more diverse and realistic perturbations across different IE tasks. Based on this dataset, we comprehensively evaluate existing UIE models and reveal that both LLM-based models and other models suffer from significant performance drops. To improve robustness and reduce training costs, we propose a data-augmentation solution that dynamically selects hard samples for iterative training based on the model's inference loss. Experimental results show that training with only \textbf{15\%} of the data leads to an average \textbf{7.5\%} relative performance improvement across three IE tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。