arXiv:2509.14399cs.CL2025-09EMNLP被引 1

用大模型重标注条件语义相似度数据,提升模型性能

Annotating Training Data for Conditional Semantic Textual Similarity Measurement using Large Language Models

  • 用大模型修正原始数据的条件描述和评分
  • 新数据训练模型,斯皮尔曼相关系数提升5.4%
  • 适合做语义相似度研究的学者和工程师

句子间的语义相似度取决于所考虑的方面。为研究这一现象,Deshpande 等人(2023)提出了条件语义文本相似度(C-STS)任务,并构建了一个包含两组不同条件下人类评分的句子对数据集。然而,Tu 等人(2024)发现该数据集存在多种标注问题,并表明手动重标注其中一小部分即可提升 C-STS 模型的准确性。尽管如此,缺乏大规模且准确标注的 C-STS 数据集仍是该任务进展的瓶颈,表现为当前模型表现不佳。为此,本文利用大语言模型(LLMs)对 Deshpande 等人(2023)提出的原始数据集中的条件陈述和相似度评分进行修正。所提方法仅需少量人工干预,即可大规模重标注用于 C-STS 任务的训练数据。更重要的是,基于清洗并重标注后的数据训练监督式 C-STS 模型,在斯皮尔曼相关性上实现 5.4% 的显著提升。重标注数据集已公开于 https://LivNLP.github.io/CSTS-reannotation。

原文摘要 · Abstract (English)

Semantic similarity between two sentences depends on the aspects considered between those sentences. To study this phenomenon, Deshpande et al. (2023) proposed the Conditional Semantic Textual Similarity (C-STS) task and annotated a human-rated similarity dataset containing pairs of sentences compared under two different conditions. However, Tu et al. (2024) found various annotation issues in this dataset and showed that manually re-annotating a small portion of it leads to more accurate C-STS models. Despite these pioneering efforts, the lack of large and accurately annotated C-STS datasets remains a blocker for making progress on this task as evidenced by the subpar performance of the C-STS models. To address this training data need, we resort to Large Language Models (LLMs) to correct the condition statements and similarity ratings in the original dataset proposed by Deshpande et al. (2023). Our proposed method is able to re-annotate a large training dataset for the C-STS task with minimal manual effort. Importantly, by training a supervised C-STS model on our cleaned and re-annotated dataset, we achieve a 5.4% statistically significant improvement in Spearman correlation. The re-annotated dataset is available at https://LivNLP.github.io/CSTS-reannotation.

语义相似度大模型应用数据标注自然语言处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。