arXiv:2512.12238cs.CLcs.AI2025-12

用多核高斯过程建模语义距离,提升文本相似度计算精度

Semantic Distance Measurement based on Multi-Kernel Gaussian Processes

  • 基于多核高斯过程构建可学习的语义距离模型
  • 在细粒度情感分类任务中表现优于传统方法
  • 适合需要自适应语义度量的自然语言处理场景

语义距离测量是计算语言学中的基础问题,用于量化文本片段之间的相似性或相关性,支撑文本检索和分类等任务。从数学角度看,语义距离可视为定义在文本空间或其表示空间上的度量。然而,大多数经典方法为固定形式,难以适应特定数据分布与任务需求。本文提出基于多核高斯过程(MK-GP)的语义距离测量方法,将文本对应的潜在语义函数建模为高斯过程,其协方差函数由马特恩与多项式核组合而成,核参数通过监督学习自动优化,而非人工设计。该方法在大语言模型的上下文学习(ICL)设置下,应用于细粒度情感分类任务,实验验证了其有效性。

原文摘要 · Abstract (English)

Semantic distance measurement is a fundamental problem in computational linguistics, providing a quantitative characterization of similarity or relatedness between text segments, and underpinning tasks such as text retrieval and text classification. From a mathematical perspective, a semantic distance can be viewed as a metric defined on a space of texts or on a representation space derived from them. However, most classical semantic distance methods are essentially fixed, making them difficult to adapt to specific data distributions and task requirements. In this paper, a semantic distance measure based on multi-kernel Gaussian processes (MK-GP) was proposed. The latent semantic function associated with texts was modeled as a Gaussian process, with its covariance function given by a combined kernel combining Matérn and polynomial components. The kernel parameters were learned automatically from data under supervision, rather than being hand-crafted. This semantic distance was instantiated and evaluated in the context of fine-grained sentiment classification with large language models under an in-context learning (ICL) setup. The experimental results demonstrated the effectiveness of the proposed measure.

语义距离高斯过程自然语言处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。