arXiv:2601.11190cs.CL2026-01中稿 · publication in Kno…被引 2

针对文档关系抽取中稀有关系样本少的问题,提出迭代优化框架提升模型表现。

DOREMI: Optimizing Long Tail Predictions in Document-Level Relation Extraction

  • 通过主动选择关键样本进行少量人工标注,迭代优化稀有关系
  • 在多个数据集上显著提升罕见关系的准确率,效果优于现有方法
  • 可兼容任意现有文档级关系抽取模型,适合资源有限的研究者

文档级关系抽取(DocRE)因依赖跨句上下文和关系类型长尾分布而面临挑战,许多关系类别训练样本极少。本文提出DOREMI——一种迭代优化框架,通过最少且精准的人工标注来增强低频关系的学习。不同于依赖大规模噪声数据或启发式去噪的方法,DOREMI主动筛选最具信息量的样本,提升训练效率与鲁棒性。该框架可适配任何现有DocRE模型,有效缓解长尾偏差,在稀有关系上的泛化能力显著增强,提供了一种可扩展的解决方案。

原文摘要 · Abstract (English)

Document-Level Relation Extraction (DocRE) presents significant challenges due to its reliance on cross-sentence context and the long-tail distribution of relation types, where many relations have scarce training examples. In this work, we introduce DOcument-level Relation Extraction optiMizing the long taIl (DOREMI), an iterative framework that enhances underrepresented relations through minimal yet targeted manual annotations. Unlike previous approaches that rely on large-scale noisy data or heuristic denoising, DOREMI actively selects the most informative examples to improve training efficiency and robustness. DOREMI can be applied to any existing DocRE model and is effective at mitigating long-tail biases, offering a scalable solution to improve generalization on rare relations.

关系抽取长尾学习文档级主动学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。