arXiv:2601.07271cs.CL2026-01Conference of the …

不依赖大模型生成数据,用实体信息实现文档级零样本关系抽取

Document-Level Zero-Shot Relation Extraction with Entity Side Information

  • 利用实体描述和上下位词等侧信息替代大模型生成数据
  • 在宏平均F1上比基线提升11.6%
  • 适合低资源语言如马来西亚英语新闻文本

文档级零样本关系抽取(DocZSRE)旨在无需针对特定关系进行训练的情况下,预测文本文档中的未见关系标签。现有方法依赖大语言模型(LLMs)生成未见标签的合成数据,但在马来英语等低资源语言中面临挑战,包括本地语言特征难以捕捉及生成内容存在事实错误的风险。本文提出基于实体侧信息的文档级零样本关系抽取框架(DocZSRE-SI),通过引入实体提及描述和实体提及上下位词等侧信息,实现不依赖LLM生成数据的零样本关系抽取。所提低复杂度模型在宏平均F1分数上相较基线模型平均提升11.6%,为高噪声、低资源场景下的关系抽取提供了稳健高效的新方案,尤其适用于马来西亚英语新闻文本等实际应用。

原文摘要 · Abstract (English)

Document-Level Zero-Shot Relation Extraction (DocZSRE) aims to predict unseen relation labels in text documents without prior training on specific relations. Existing approaches rely on Large Language Models (LLMs) to generate synthetic data for unseen labels, which poses challenges for low-resource languages like Malaysian English. These challenges include the incorporation of local linguistic nuances and the risk of factual inaccuracies in LLM-generated data. This paper introduces Document-Level Zero-Shot Relation Extraction with Entity Side Information (DocZSRE-SI) to address limitations in the existing DocZSRE approach. The DocZSRE-SI framework leverages Entity Side Information, such as Entity Mention Descriptions and Entity Mention Hypernyms, to perform ZSRE without depending on LLM-generated synthetic data. The proposed low-complexity model achieves an average improvement of 11.6% in the macro F1-Score compared to baseline models and existing benchmarks. By utilizing Entity Side Information, DocZSRE-SI offers a robust and efficient alternative to error-prone, LLM-based methods, demonstrating significant advancements in handling low-resource languages and linguistic diversity in relation extraction tasks. This research provides a scalable and reliable solution for ZSRE, particularly in contexts like Malaysian English news articles, where traditional LLM-based approaches fall short.

零样本学习关系抽取低资源语言实体信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。