arXiv:2506.21222cs.CLcs.IR2025-06ACL被引 4

用句法相似性选示例,提升大模型术语提取准确率

Enhancing Automatic Term Extraction with Large Language Models via Syntactic Retrieval

  • 根据句法结构而非语义选相似示例
  • 在三个专业数据集上F1分数均提升
  • 适合需要精准术语识别的领域应用

自动术语提取(ATE)能识别对机器翻译、信息检索等任务至关重要的领域特有表达。尽管大语言模型(LLMs)已显著推动诸多NLP任务发展,但其在ATE中的潜力尚未得到充分探索。本文提出一种基于检索的提示策略,在少样本设置下,依据句法而非语义相似性选择示范样本。该句法检索方法具有领域无关性,能更可靠地帮助捕捉术语边界。我们在同领域和跨领域设置下进行评估,分析查询句与检索示例间的词汇重叠对性能的影响。在三个专用ATE基准上的实验表明,句法检索提升了F1分数。这些发现强调了在将LLMs应用于术语提取任务时,句法线索的重要性。

原文摘要 · Abstract (English)

Automatic Term Extraction (ATE) identifies domain-specific expressions that are crucial for downstream tasks such as machine translation and information retrieval. Although large language models (LLMs) have significantly advanced various NLP tasks, their potential for ATE has scarcely been examined. We propose a retrieval-based prompting strategy that, in the few-shot setting, selects demonstrations according to \emph{syntactic} rather than semantic similarity. This syntactic retrieval method is domain-agnostic and provides more reliable guidance for capturing term boundaries. We evaluate the approach in both in-domain and cross-domain settings, analyzing how lexical overlap between the query sentence and its retrieved examples affects performance. Experiments on three specialized ATE benchmarks show that syntactic retrieval improves F1-score. These findings highlight the importance of syntactic cues when adapting LLMs to terminology-extraction tasks.

术语提取大模型句法特征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。