用模拟人工标注分歧来优化指令,提升零样本实体识别效果
DiZiNER: Disagreement-guided Instruction Refinement via Pilot Annotation Simulation for Zero-shot Named Entity Recognition

- 让多个大模型互相标注文本,通过分析分歧改进任务指令
- 在18个数据集上14个达顶尖水平,性能提升8.0点F1
- 适合想提升零样本命名实体识别效果的研究者
大语言模型虽推动了零样本和少样本命名实体识别发展,但生成结果仍存在系统性错误。尽管经过指令微调,零样本表现仍远低于有监督模型。这些错误类似早期人工标注中的不一致问题,可通过预标注流程解决。受此启发,我们提出DiZiNER框架,模拟预标注过程,让多个异构大模型作为标注者与监督者协同工作。多个大模型对共享文本进行标注,由监督模型分析模型间分歧,进而优化任务指令。在18个基准上,DiZiNER在14个数据集上达到零样本最优(SOTA),F1提升最高达+8.0,零样本与有监督差距缩小超过+11点。其表现优于监督模型GPT-5 mini,说明提升来自分歧引导的指令优化而非模型能力。模型间配对一致性与NER性能呈强相关,进一步支持该结论。
原文摘要 · Abstract (English)
Large language models (LLMs) have advanced information extraction (IE) by enabling zero-shot and few-shot named entity recognition (NER), yet their generative outputs still show persistent and systematic errors. Despite progress through instruction fine-tuning, zero-shot NER still lags far behind supervised systems. These recurring errors mirror inconsistencies observed in early-stage human annotation processes that resolve disagreements through pilot annotation. Motivated by this analogy, we introduce DiZiNER (Disagreement-guided Instruction Refinement via Pilot Annotation Simulation for Zero-shot Named Entity Recognition), a framework that simulates the pilot annotation process, employing LLMs to act as both annotators and supervisors. Multiple heterogeneous LLMs annotate shared texts, and a supervisor model analyzes inter-model disagreements to refine task instructions. Across 18 benchmarks, DiZiNER achieves zero-shot SOTA results on 14 datasets, improving prior bests by +8.0 F1 and reducing the zero-shot to supervised gap by over +11 points. It also consistently outperforms its supervisor, GPT-5 mini, indicating that improvements stem from disagreement-guided instruction refinement rather than model capacity. Pairwise agreement between models shows a strong correlation with NER performance, further supporting this finding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。