arXiv:2506.18602cs.CLstat.AP2025-06中稿 · publication in the…被引 2

BERT在特定领域语义相似度计算中表现最优。

Semantic similarity estimation for domain specific data using BERT and other techniques

  • 对比USE、InferSent与BERT,采用微调提升语义理解能力。
  • 在两个问答对数据集上,BERT显著优于其他方法。
  • 适合需要高精度语义匹配的垂直领域应用。

语义相似度估计是自然语言处理与理解中的重要问题,在问答系统、语义搜索、信息检索、文档聚类、词义消歧和机器翻译等下游任务中具有广泛应用。本文对比了USE(Universal Sentence Encoder)、InferSent及最新的BERT模型在语义相似度估计上的表现。实验使用两个问答对数据集:一个特定领域的内部数据集和公开的Quora问答对数据集。结果表明,经过微调的BERT模型性能显著优于其他方法,这得益于其训练过程中基于数据的学习能力。该研究验证了BERT在特定领域数据上的适用性,结论认为其是处理领域特定数据的最佳选择。

原文摘要 · Abstract (English)

Estimation of semantic similarity is an important research problem both in natural language processing and the natural language understanding, and that has tremendous application on various downstream tasks such as question answering, semantic search, information retrieval, document clustering, word-sense disambiguation and machine translation. In this work, we carry out the estimation of semantic similarity using different state-of-the-art techniques including the USE (Universal Sentence Encoder), InferSent and the most recent BERT, or Bidirectional Encoder Representations from Transformers, models. We use two question pairs datasets for the analysis, one is a domain specific in-house dataset and the other is a public dataset which is the Quora's question pairs dataset. We observe that the BERT model gave much superior performance as compared to the other methods. This should be because of the fine-tuning procedure that is involved in its training process, allowing it to learn patterns based on the training data that is used. This works demonstrates the applicability of BERT on domain specific datasets. We infer from the analysis that BERT is the best technique to use in the case of domain specific data.

语义相似度BERT领域适应文本匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。