arXiv:2509.11513cs.CLcs.AI2025-09

用上下文影响分析提升词汇替换候选排序准确率

Unsupervised Candidate Ranking for Lexical Substitution via Holistic Sentence Semantics

  • 基于注意力与梯度方法衡量上下文对目标词的影响
  • 在LS07和SWORDS数据集上提升排序性能
  • 无需调参,适合需要可解释性的自然语言任务

词汇替换中的关键子任务是候选词排序。传统方法通过将目标词替换为候选词后输入模型,捕捉替换前后的语义差异。然而,有效建模候选词对目标词及其上下文的双向影响仍具挑战。现有方法多仅关注目标位置的语义变化,或依赖多指标参数调优,难以准确刻画语义变化。为此,本文提出两种方法:一种基于注意力权重,另一种采用更可解释的积分梯度法,均用于测量上下文对目标词的影响,并结合原句与替换句的语义相似度进行候选词排序。在LS07和SWORDS数据集上的实验表明,两种方法均提升了排序性能。

原文摘要 · Abstract (English)

A key subtask in lexical substitution is ranking the given candidate words. A common approach is to replace the target word with a candidate in the original sentence and feed the modified sentence into a model to capture semantic differences before and after substitution. However, effectively modeling the bidirectional influence of candidate substitution on both the target word and its context remains challenging. Existing methods often focus solely on semantic changes at the target position or rely on parameter tuning over multiple evaluation metrics, making it difficult to accurately characterize semantic variation. To address this, we investigate two approaches: one based on attention weights and another leveraging the more interpretable integrated gradients method, both designed to measure the influence of context tokens on the target token and to rank candidates by incorporating semantic similarity between the original and substituted sentences. Experiments on the LS07 and SWORDS datasets demonstrate that both approaches improve ranking performance.

词汇替换语义排序可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。