用自修正迭代框架提升信息抽取效率与准确率
SCIR: A Self-Correcting Iterative Refinement Framework for Enhanced Information Extraction Based on Schema
- 通过双路径自修正模块实现零成本接入现有大模型
- 在三项任务上平均提升5.27%准确率,训练成本降低87%
- 适合追求高效低耗信息抽取的开发者与研究者
尽管基于大语言模型(LLM)的信息抽取(IE)系统展现出强大能力,但当前微调范式存在训练成本高和与大模型偏好难以对齐两大问题。为此,我们提出一种通用的IE新范式——自修正迭代精炼(SCIR)框架,并构建包含超10万条数据的多任务双语(中英)自修正(MBSC)数据集。SCIR框架通过双路径自修正模块与反馈驱动优化,实现与现有LLM和IE系统的即插即用兼容,显著降低训练成本。同时,MBSC数据集通过间接蒸馏GPT-4的能力,解决偏好对齐难题。实验表明,SCIR在命名实体识别、关系抽取和事件抽取三个关键任务上均优于现有方法,跨度级微观F1平均提升5.27%,训练成本较基线降低87%。该成果不仅提升了IE系统的灵活性与准确性,也为轻量高效的IE范式发展提供新路径。
原文摘要 · Abstract (English)
Although Large language Model (LLM)-powered information extraction (IE) systems have shown impressive capabilities, current fine-tuning paradigms face two major limitations: high training costs and difficulties in aligning with LLM preferences. To address these issues, we propose a novel universal IE paradigm, the Self-Correcting Iterative Refinement (SCIR) framework, along with a Multi-task Bilingual (Chinese-English) Self-Correcting (MBSC) dataset containing over 100,000 entries. The SCIR framework achieves plug-and-play compatibility with existing LLMs and IE systems through its Dual-Path Self-Correcting module and feedback-driven optimization, thereby significantly reducing training costs. Concurrently, the MBSC dataset tackles the challenge of preference alignment by indirectly distilling GPT-4's capabilities into IE result detection models. Experimental results demonstrate that SCIR outperforms state-of-the-art IE methods across three key tasks: named entity recognition, relation extraction, and event extraction, achieving a 5.27 percent average improvement in span-based Micro-F1 while reducing training costs by 87 percent compared to baseline approaches. These advancements not only enhance the flexibility and accuracy of IE systems but also pave the way for lightweight and efficient IE paradigms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。