arXiv:2411.11682stat.MLcs.LG2024-11被引 1

用对比学习训练可微分的结构化损失,提升复杂输出预测效果

Learning Differentiable Surrogate Losses for Structured Prediction

  • 通过对比学习从数据中直接学习神经网络参数化的损失函数
  • 在图预测任务上表现优于或媲美预定义核函数的方法
  • 适合需要自适应设计损失的复杂结构预测场景

结构化预测旨在预测复杂结构而非简单标量。其主要挑战源于输出空间的非欧几里得特性,通常需放松问题形式化。代理方法基于核诱导损失或更一般的具有隐式损失嵌入的损失函数,将原问题转化为回归任务再经解码步骤求解。然而,为复杂结构设计有效损失函数面临巨大挑战,且常需领域专业知识。本文提出新框架:在处理监督代理回归前,通过对比学习从输出训练数据中直接学习由神经网络参数化的结构化损失函数。结果,该可微分损失不仅因代理空间有限维使神经网络可训练,还支持基于梯度下降的解码策略来预测新输出结构。数值实验显示,在监督图预测任务中,本方法性能与甚至优于基于预定义核的方法。

原文摘要 · Abstract (English)

Structured prediction involves learning to predict complex structures rather than simple scalar values. The main challenge arises from the non-Euclidean nature of the output space, which generally requires relaxing the problem formulation. Surrogate methods build on kernel-induced losses or more generally, loss functions admitting an Implicit Loss Embedding, and convert the original problem into a regression task followed by a decoding step. However, designing effective losses for objects with complex structures presents significant challenges and often requires domain-specific expertise. In this work, we introduce a novel framework in which a structured loss function, parameterized by neural networks, is learned directly from output training data through Contrastive Learning, prior to addressing the supervised surrogate regression problem. As a result, the differentiable loss not only enables the learning of neural networks due to the finite dimension of the surrogate space but also allows for the prediction of new structures of the output data via a decoding strategy based on gradient descent. Numerical experiments on supervised graph prediction problems show that our approach achieves similar or even better performance than methods based on a pre-defined kernel.

结构化预测对比学习可微分损失图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。