arXiv:2510.10241cs.CLcs.IR2025-10中稿 · ACL2026 main被引 1

用轻量模型+大模型纠错,提升指代消解准确率

ImCoref-CeS: An Improved Lightweight Pipeline for Coreference Resolution with LLM-based Checker-Splitter Refinement

  • 轻量模型增强长文本编码与位置感知能力
  • 大模型充当检查员和拆分器,修正错误指代结果
  • 适合追求高精度且资源受限的指代消解场景

指代消解(CR)是自然语言处理中的关键任务。当前研究面临抉择:是继续挖掘小语言模型在监督神经方法上的潜力,其检测-聚类流程仍保持领先性能,还是拥抱大语言模型(LLMs)的强大能力?但两者优势的有效结合仍待探索。为此,我们提出新框架ImCoref-CeS,融合改进的监督模型与基于LLM的推理。首先,我们设计了改进的CR方法(ImCoref),通过引入轻量桥接模块增强长文本编码能力,采用双仿射评分器全面捕捉位置信息,并使用混合提及正则化提升训练效率。更重要的是,我们让一个大模型扮演多角色检查员-拆分器,验证候选提及(过滤无效项)并修正ImCoref的指代结果(拆分错误聚类)。大量实验表明,ImCoref-CeS在性能上优于现有最先进方法。

原文摘要 · Abstract (English)

Coreference Resolution (CR) is a critical task in Natural Language Processing (NLP). Current research faces a key dilemma: whether to further explore the potential of supervised neural methods based on small language models, whose detect-then-cluster pipeline still delivers top performance, or embrace the powerful capabilities of Large Language Models (LLMs). However, effectively combining their strengths remains underexplored. To this end, we propose \textbf{ImCoref-CeS}, a novel framework that integrates an enhanced supervised model with LLM-based reasoning. First, we present an improved CR method (\textbf{ImCoref}) to push the performance boundaries of the supervised neural method by introducing a lightweight bridging module to enhance long-text encoding capability, devising a biaffine scorer to comprehensively capture positional information, and invoking a hybrid mention regularization to improve training efficiency. Importantly, we employ an LLM acting as a multi-role Checker-Splitter agent to validate candidate mentions (filtering out invalid ones) and coreference results (splitting erroneous clusters) predicted by ImCoref. Extensive experiments demonstrate the effectiveness of ImCoref-CeS, which achieves superior performance compared to existing state-of-the-art (SOTA) methods.

指代消解大模型应用轻量模型纠错机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。