arXiv:2608.08422stat.MEcs.LG2026-08

提出生成真实排名数据的新框架,解决隐私保护与模拟难题。

Population-Level Generative Modeling for Ranking Data

论文配图:Population-Level Generative Modeling for Ranking Data
图 1 · 摘自论文原文
  • 用低维偏好单纯形嵌入建模群体偏好分布
  • 通过流匹配学习潜变量分布,生成高保真排名数据
  • 适合需要模拟用户偏好的推荐系统与评测场景

排名数据广泛存在于推荐系统、信息检索、投票、营销及基于人类反馈的AI偏好排序等场景。现有统计方法多聚焦于推断任务,如偏好估计、排名聚合与预测,但缺乏从观测群体中生成真实合成排名的能力,而这一能力对隐私保护数据共享、基准构建、模拟和不确定性量化至关重要。该任务挑战在于排名是高维组合对象,具有非欧几里得依赖结构,且群体偏好异质性强。本文提出一种基于潜在偏好单纯形嵌入的人群级生成建模框架:通过似然型排名模型估计低维潜变量偏好单纯形,利用流匹配学习潜变量在群体中的分布,并通过拟合的概率排名模型生成新排名。我们证明排名生成可归约为潜变量分布学习的“预言机”问题,推导出有限样本生成保证,阐明项目数、排名长度与潜维数如何影响生成精度。在合成与真实数据集上的实验表明,该方法显著提升人群级保真度,并提供可解释的偏好异质性表示。

原文摘要 · Abstract (English)

Ranking data arise in scientific and machine learning applications, including recommendation systems, information retrieval, voting, marketing, and AI preference ranking from human feedback. Existing statistical work has primarily focused on inference tasks such as preference estimation, rank aggregation, and ranking prediction. However, generating realistic synthetic rankings from an observed population is important for privacy-preserving data sharing, benchmark construction, simulation, and uncertainty quantification. This task is challenging because rankings are high-dimensional combinatorial objects with non-Euclidean dependence structures, while ranking populations often exhibit substantial preference heterogeneity. We propose a framework for population-level generative modeling through a latent preference simplex embedding. It estimates a low-dimensional latent preference simplex through a likelihood-based ranking model, leverages flow matching to learn the population distribution of latent preferences, and generates new rankings through the fitted probabilistic ranking model. We show that ranking generation admits an oracle reduction to latent distribution learning and derive finite-sample generative guarantees that clarify how the number of items, ranking length, and latent dimension affect accuracy. Experiments on synthetic and real datasets demonstrate improved population-level fidelity and provide a statistically interpretable representation of preference heterogeneity.

生成模型排名数据潜变量建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。