arXiv:2605.27619cs.LGcs.AI2026-05

用最优传输与依赖最大化,让模型同时学好数据结构和预测信号。

Supervised Distributional Reduction via Optimal Transport and Dependence Maximization

论文配图:Supervised Distributional Reduction via Optimal Transport and Dependence Maximization
图 1 · 摘自论文原文
  • 结合最优传输与显式依赖项,学习目标感知的紧凑表示
  • 在保留几何结构的同时,显著提升下游预测性能
  • 适合需要融合结构与标签信息的场景,如高维数据建模

学习既能捕捉数据内在几何结构又能保留目标相关特征的表示,仍是核心挑战,尤其在压缩与预测保真度之间需权衡时。分布式降维(涵盖聚类与降维)提供了一种原则性方法来总结数据,但其有监督版本仍研究不足。本文提出有监督分布式降维(SDR),通过最优传输与显式依赖最大化联合优化,构建输入分布与代表性点之间的关系对齐,并引入直接依赖项以更明确地捕获预测信号。所得表示兼具几何结构与监督信息。此外,SDR自然生成依赖数据的非平稳几何,可用于高斯过程建模。通过目标感知的分布对齐重新定义距离,可构造响应局部数据结构与监督变化的自适应核函数,为非平稳核设计提供了基于最优传输的新视角。

原文摘要 · Abstract (English)

Learning representations that capture both intrinsic data geometry and target-relevant structure remains a fundamental challenge, particularly in settings where data reduction must balance compression with predictive fidelity. While distributional reduction-encompassing joint clustering and dimensionality reduction-offers a principled way to summarize data, its supervised variants remain relatively under-explored, despite the importance of retaining task-relevant signal for downstream prediction and decision-making. We propose Supervised Distributional Reduction (SDR), an algorithm for learning target-aware representations by combining optimal transport with explicit dependence maximization. SDR builds on the Fused Gromov-Wasserstein (FGW) objective to align the relational structure of the input distribution with a set of representative points, while augmenting it with a direct dependence term that encourages the learned embeddings to capture predictive signal more explicitly. This results in compact representations that reflect both geometric structure and supervision. Beyond representation learning, SDR naturally induces a data-dependent, non-stationary geometry that can be leveraged for settings such as Gaussian Process (GP) modelling. By redefining distances through target-aware distributional alignment, SDR enables the construction of adaptive kernels that respond to local variations in both data geometry and supervision, offering an optimal transport-based perspective on non-stationary kernel design.

表示学习最优传输监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。