arXiv:2512.18056cs.LGcs.SI2025-12被引 1

用概率模型构建可解释的用户数字孪生,揭示行为背后的稳定特质。

Probabilistic Digital Twins of Users: Latent Representation Learning with Statistically Validated Semantics

  • 将用户建模为生成行为数据的随机潜状态,通过变分推断学习。
  • 发现用户特征主要呈连续分布,少数主导轴线对应意见强度等可解释特质。
  • 结合非参数检验验证潜空间语义,适合需可解释推荐与个性化系统的研究者。

理解用户身份与行为对个性化、推荐和决策支持至关重要。现有方法多依赖确定性嵌入或黑箱模型,缺乏不确定性量化且难以解释潜表示内涵。本文提出一种概率数字孪生框架,将每位用户视为生成观测行为数据的潜在随机状态,通过摊销变分推断实现可扩展的后验估计,并保持全概率解释。在设计用于捕捉用户身份稳定性的用户响应数据集上,应用变分自编码器实例化该框架。除标准重建评估外,引入基于统计的解释管道,分析各潜维度极端用户的可观测行为模式,通过非参数假设检验与效应量验证,证明特定维度对应意见强度、决断力等可解释特质。实证表明,用户结构主要呈连续分布而非离散聚类,仅少数主导潜轴表现出弱但有意义的结构。结果表明,概率数字孪生可提供可解释且含不确定性的用户表征,超越传统确定性嵌入。

原文摘要 · Abstract (English)

Understanding user identity and behavior is central to applications such as personalization, recommendation, and decision support. Most existing approaches rely on deterministic embeddings or black-box predictive models, offering limited uncertainty quantification and little insight into what latent representations encode. We propose a probabilistic digital twin framework in which each user is modeled as a latent stochastic state that generates observed behavioral data. The digital twin is learned via amortized variational inference, enabling scalable posterior estimation while retaining a fully probabilistic interpretation. We instantiate this framework using a variational autoencoder (VAE) applied to a user-response dataset designed to capture stable aspects of user identity. Beyond standard reconstruction-based evaluation, we introduce a statistically grounded interpretation pipeline that links latent dimensions to observable behavioral patterns. By analyzing users at the extremes of each latent dimension and validating differences using nonparametric hypothesis tests and effect sizes, we demonstrate that specific dimensions correspond to interpretable traits such as opinion strength and decisiveness. Empirically, we find that user structure is predominantly continuous rather than discretely clustered, with weak but meaningful structure emerging along a small number of dominant latent axes. These results suggest that probabilistic digital twins can provide interpretable, uncertainty-aware representations that go beyond deterministic user embeddings.

用户建模概率推理可解释性潜变量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。