用大模型生成解释,提升中文低资源命名实体识别效果
Improving Low-Resource Sequence Labeling with Knowledge Fusion and Contextual Label Explanations
- 用大模型生成实体上下文解释,缓解领域语义偏差
- 新模型能高效提取嵌套实体,推理时无需外部知识
- 在多个中文专业数据集上达到最优性能,适合低资源场景
序列标注在低资源、领域特定场景下仍具挑战性,尤其对汉字密集型语言如中文。现有方法多关注提升模型理解力和数据多样性,但仍存在模型适用性不足及领域内语义分布偏差问题。为此,我们提出一个结合大模型知识增强流程与基于跨度的KnowFREE(Knowledge Fusion for Rich and Efficient Extraction)模型的新框架。该流程通过解释提示生成目标实体的精准上下文解释,有效缓解语义偏差并丰富模型上下文理解。KnowFREE模型进一步融合扩展标签特征,实现无需外部知识即可高效提取嵌套实体。在多个中文领域特定序列标注数据集上的实验表明,该方法性能达当前最优,有效应对低资源场景挑战。
原文摘要 · Abstract (English)
Sequence labeling remains a significant challenge in low-resource, domain-specific scenarios, particularly for character-dense languages like Chinese. Existing methods primarily focus on enhancing model comprehension and improving data diversity to boost performance. However, these approaches still struggle with inadequate model applicability and semantic distribution biases in domain-specific contexts. To overcome these limitations, we propose a novel framework that combines an LLM-based knowledge enhancement workflow with a span-based Knowledge Fusion for Rich and Efficient Extraction (KnowFREE) model. Our workflow employs explanation prompts to generate precise contextual interpretations of target entities, effectively mitigating semantic biases and enriching the model's contextual understanding. The KnowFREE model further integrates extension label features, enabling efficient nested entity extraction without relying on external knowledge during inference. Experiments on multiple Chinese domain-specific sequence labeling datasets demonstrate that our approach achieves state-of-the-art performance, effectively addressing the challenges posed by low-resource settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。